Summary We're excited to introduce Provisioned Throughput, reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA. Open weight models have become the essential ingredient for any company that wants an abundant AI future. But historically to use open weight models companies have had to choose between convenient but best-effort serverless or powerful but tunable dedicated inference. Provisioned Throughput is a new inference form factor that offers an alternate option that gives customers the simplicity of pre-optimized models at token prices with the predictability of guaranteed capacity in with an availability SLA. Costs run up to 90% below Claude Opus 4.8 at list price. Available today for MiniMax M3 and GLM-5.2, with capacity in North America, EMEA, and beyond, and a one-month minimum term. Inference spend has become a line item the board asks about. In a not too distant future every knowledge worker will have several agents working on their behalf, but companies are struggling to achieve this without astronomic inference budgets. …