Summary We're releasing a significant update to our inference platform, giving you complete control over performance, cost, and quality without building your own stack. Models go live in minutes and every deployment is production-grade from the start: run multiple deployments behind one stable endpoint, ship changes safely with canary, blue-green, and rolling updates that auto-roll-back on your thresholds, test on real traffic with A/B and shadow testing, and autoscale across one region or many. We're also announcing a closed beta for custom training, including full-weight and LoRA reinforcement learning and supervised fine-tuning, with checkpoints you can deploy straight to production. Request access to the custom training beta. Launch your first endpoint today. Join our webinar on August 6 for a deep dive. Open-weight models now match closed models on quality, run at a fraction of the cost, and are fully customizable for your task. That combination has made them the foundation for teams building serious AI products and agents. But the strategic motivation for adopting open-weight models is still the ability to exercise control. …