A new comparison of DeepSeek V4 Pro 0813 and GPT-5.6 Sol shows that a cascading strategy—using DeepSeek V4 Pro 0813 first and escalating to GPT-5.6 Sol when tests fail—solves 83.0% of DeepSWE tasks…
Together AI compares DeepSeek V4 Pro 0813 and Claude Fable 5 on DeepSWE, finding that a cascade strategy (Pro first, escalate to Fable) solves 82.7% of tasks at $8.28 each, outperforming Fable alone…
Together AI introduces native A/B testing for LLM endpoints, allowing traffic splits between a control and up to 20 variants with fixed percentages, etag-guarded updates, and blue-green promotion.
Key Takeaways While GPT-5.6 Luna is the stronger engineer on every quality measure, DeepSeek-V4 Flash 0731 is cheap enough that a DeepSeek-first cascade beats Luna alone on both accuracy and cost.
What you'll learn What is Kimi K3, and what makes it different? What is under the hood: KDA, Attention Residuals, and the Stable LatentMoE architecture How do you use reasoning effort, streaming,…
Summary With Dedicated Model Inference on the Together AI platform y ou can get your deployments to autoscale on metrics the inference engine actually understands, such as in-flight requests, TTFT,…
Summary We introduce ThunderAgent, a system for high throughput agentic inference. By introducing a novel program abstraction for agentic LLM request scheduling, ThunderAgent achieves up to 2.5×…
Summary Dedicated Model Inference on the Together AI platform consists of three parts: the endpoint (a stable name you or your clients call), deployments (specific model + hardware combinations…
Today we're announcing a strategic partnership with Moonshot AI, one of the leading model labs pushing the open source frontier with large-scale MoE architectures.
Key Takeaways Kimi K3 matches Claude Fable 5 on quality, costs a third as much per solved task, and as an open model gives teams full control over their deployment.
Summary We're releasing a significant update to our inference platform, giving you complete control over performance, cost, and quality without building your own stack.
Today, Together AI and Y Combinator (YC) are announcing a partnership to deliver the first dedicated YC GPU cluster, giving YC's portfolio of AI-native startups easier access to the compute they need…
We’ve spent the last several weeks shipping a set of changes to Together GPU Clusters aimed at the operational reality of running training and inference at scale: hardware fails, schedulers leak, and…
Today, Thinking Machines Lab released Inkling, a new multimodal mixture-of-experts model built for token-efficient reasoning, native multimodal understanding, and broad task versatility.
Summary We're excited to introduce Provisioned Throughput, reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA.
Summary LLMs have gotten surprisingly good at writing GPU kernels [1][2][3] , but almost all current benchmarks measuring that progress are single-GPU.
Summary Together AI has received an ISO 27001:2022 certification from A-LIGN Compliance and Security, Inc., an ANAB-accredited certification body, confirming that our Information Security Management…
Together AI is the preferred cloud partner for MiniMax M3. Together AI will host the open-weights model as a developer endpoint upon its public release.
Artificial Analysis reported speed factor (Input audio seconds transcribed per second) -- Higher is better Modality matters A 1M-token text prompt can fit the entire Harry Potter series and still…
Summary On a production coding agent workload, Together Inference Engine delivers 31% more TPS than the next fastest OSS engine on the same hardware, and maintains 2× better TTFT at saturation.
Together AI’s mission is to make AI accessible and widely distributed. Today, we are excited to announce an exclusive partnership with Pearl Research Labs, the creators of the Pearl Network, a…
Video has become one of the most popular mediums for information sharing. Yet, the language distribution of popular video contents on the internet does not necessarily reflect the diversity of global…
Benchmark tables miss the main point of DeepSeek-V4: the important change is architectural. V4 turns million-token context into a serving-systems problem.
Something real is shifting in how developers work. Agents open up work that used to be off-limits, not because it was technically impossible, but because it required niche expertise most of us didn't…
We’re excited to partner with Adaption to make Together Fine-Tuning available in Adaptive Data . Adaption is co-founded by Sara Hooker and Sudip Roy, both former leaders at Cohere and Google DeepMind…
Summary We were able to get ahead of Copy Fail (CVE‑2026‑31431) by treating it as a fleet‑level emergency, shutting off the vulnerable crypto socket interface across our infrastructure within hours…
What's New DeepSeek V4 Pro on Together AI: DeepSeek V4 Pro is now available on Together AI with a 512K-token context window for long-context reasoning workloads.
NVIDIA Nemotron™ 3 Nano Omni is now available on the Together AI platform. Representing a meaningful step forward for multimodal AI, Nemotron 3 Nano Omni is a single, open model that reasons across…
Summary Distribution-aware speculative decoding (DAS) is a novel framework that significantly alleviates the rollout bottleneck in RL post-training — delivering up to 50% speedup without touching…
Summary We present Parcae, one of the first stable architectures for looped language models , achieving the quality of a Transformer twice the size with clean, predictable training.
Summary Scientific discovery has driven human progress, but tackling today’s hardest problems requires collective intelligence beyond any single researcher or model.
Over the last few years of powering and partnering with the fastest scaling AI-native companies, we have come to realize they need a different kind of cloud: an AI Native Cloud.
Summary We worked in collaboration with Stanford University, the University of Wisconsin–Madison, and Bauplan to test whether LLMs can optimize database query execution plans.
Summary Four-model suite: Wan 2.7 brings video generation, continuation, and editing to Together AI, starting today with text-to-video and expanding soon to image-to-video, reference-to-video, and…
Deepgram Nova-3 , Nova-3 Multilingual , Flux , and Aura-2 now run natively on Together AI Dedicated Model Inference Deepgram covers both ends of the voice pipeline, from transcription to synthesis,…
The breakthrough came on a holiday weekend. Memorial Day 2022. While most of Silicon Valley was at barbecues, Dan Fu, Tri Dao, and their colleagues were about to prove the AI establishment wrong.
What’s New Tool call fine-tuning: Ensure agents execute structured actions reliably with end-to-end fine-tuning and inference on OpenAI-compatible schema.
tl;dr Mamba-3 is a new state space model (SSM) designed with inference efficiency as the primary goal — a departure from Mamba-2, which optimized for training speed.
This year, Together AI is excited to be part of NVIDIA GTC with multiple major announcements and conversations shaping the AI ecosystem — from cutting-edge model releases to new voice AI…
Summary Together AI, the AI Native Cloud , announced a full suite of capabilities for building real-time voice agents — co-located STT, LLM, and TTS on one cloud, eliminating inter-vendor network…