Anthropic is betting that a purpose-built inference chip can cut serving cost per token enough to justify the switching risk.
Order volume disclosed to hardware partners, and detailed in a report from The Information, implies the new chips will carry a majority of production inference traffic within two quarters.
The economics behind the bet
Inference, not training, is now the larger and stickier cost line for a model lab at Anthropic's scale — every serving request has a marginal cost that compounds across billions of daily calls, in a way a one-time training run does not.
The move mirrors a broader pattern: model labs that once treated compute as a pure procurement problem are now treating it as a design constraint on the model itself.
Custom silicon lets Anthropic co-design model architecture and chip layout in tandem, something that's structurally hard to do when renting general-purpose GPUs from a cloud provider.
The skeptics' case
Skeptics note that custom silicon has a poor track record outside the hyperscalers that design their own chips end-to-end — a model lab going the same route inherits all of the integration risk with none of the decade of tooling investment those hyperscalers have already sunk into their own stacks.
A Bloomberg analysis of prior custom-silicon efforts at other AI labs found mixed results, with several quietly reverting to merchant GPUs after yield or software-tooling problems.
— Marcus Reyes
