← FeedLeaderboard →

Together AI

@togethercompute

Last 8 weeks · 19 posts · 0 outlets writing · no. 10 of 20 by coverage

Window4w8w12w26w

Share of voice, week by week

29 Jun–5 Jul→ 17–23 Aug…
29 Jun–5 Jul6–12 Jul13–19 Jul20–26 Jul27 Jul–2 Aug3–9 Aug10–16 Aug17–23 Aug…
Postspeak 5/wk
Coveragepeak 1/wk

Who covers them

No outlet coverage resolved inside this window.

Recent stories

Together AIModel update13h ago

DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Together AIModel update· 13h ago

A new comparison of DeepSeek V4 Pro 0813 and GPT-5.6 Sol shows that a cascading strategy—using DeepSeek V4 Pro 0813 first and escalating to GPT-5.6 Sol when tests fail—solves 83.0% of DeepSWE tasks…

Together AIModel update1 day ago

DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

Together AIModel update· 1 day ago

Together AI compares DeepSeek V4 Pro 0813 and Claude Fable 5 on DeepSWE, finding that a cascade strategy (Pro first, escalate to Fable) solves 82.7% of tasks at $8.28 each, outperforming Fable alone…

Together AIProduct launch1 day ago

A/B test models in production

Together AIProduct launch· 1 day ago

Together AI introduces native A/B testing for LLM endpoints, allowing traffic splits between a control and up to 20 variants with fixed percentages, etag-guarded updates, and blue-green promotion.

Together AI12 days ago

DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

Together AI· 12 days ago

Key Takeaways While GPT-5.6 Luna is the stronger engineer on every quality measure, DeepSeek-V4 Flash 0731 is cheap enough that a DeepSeek-first cascade beats Luna alone on both accuracy and cost.

Together AI17 days ago

Kimi K3: The Complete Developer Guide

Together AI· 17 days ago

What you'll learn What is Kimi K3, and what makes it different? What is under the hood: KDA, Attention Residuals, and the Stable LatentMoE architecture How do you use reasoning effort, streaming,…

Together AI18 days ago

Autoscaling endpoints for LLM inference

Together AI· 18 days ago

Summary With Dedicated Model Inference on the Together AI platform y ou can get your deployments to autoscale on metrics the inference engine actually understands, such as in-flight requests, TTFT,…

Together AI20 days ago

ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

Together AI· 20 days ago

Summary We introduce ThunderAgent, a system for high throughput agentic inference. By introducing a novel program abstraction for agentic LLM request scheduling, ThunderAgent achieves up to 2.5×…

Together AI20 days ago

Configuring Dedicated Model Inference

Together AI· 20 days ago

Summary Dedicated Model Inference on the Together AI platform consists of three parts: the endpoint (a stable name you or your clients call), deployments (specific model + hardware combinations…

Together AI20 days ago

Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models

Together AI· 20 days ago

Today we're announcing a strategic partnership with Moonshot AI, one of the leading model labs pushing the open source frontier with large-scale MoE architectures.

Together AI23 days ago

Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Together AI· 23 days ago

Summary GPT-5.6 Sol edges Kimi K3 on single-shot quality, but Kimi wins on pass@k with k > 1 and costs 64% less per completed task.

Together AI25 days ago

Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

Together AI· 25 days ago

Key Takeaways Kimi K3 matches Claude Fable 5 on quality, costs a third as much per solved task, and as an open model gives teams full control over their deployment.

Together AI26 days ago

The production platform for open-weight AI inference

Together AI· 26 days ago

Summary We're releasing a significant update to our inference platform, giving you complete control over performance, cost, and quality without building your own stack.

Together AI29 days ago

Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

Together AI· 29 days ago

Today, Together AI and Y Combinator (YC) are announcing a partnership to deliver the first dedicated YC GPU cluster, giving YC's portfolio of AI-native startups easier access to the compute they need…

Together AI16 Jul

What does 99.9% uptime mean for inference?

Together AI· 16 Jul

TL;DR The short version: Each reliability tier maps to a specific failure domain, and each one requires its own architecture to survive it.

Together AI15 Jul

New in Together GPU Clusters: Reliability and control for production GPU clusters

Together AI· 15 Jul

We’ve spent the last several weeks shipping a set of changes to Together GPU Clusters aimed at the operational reality of running training and inference at scale: hardware fails, schedulers leak, and…

Together AI15 Jul

Together AI brings Thinking Machines Lab’s new model Inkling on day 0

Together AI· 15 Jul

Today, Thinking Machines Lab released Inkling, a new multimodal mixture-of-experts model built for token-efficient reasoning, native multimodal understanding, and broad task versatility.

Together AI8 Jul

Open, convenient and predictable: Introducing Provisioned Throughput

Together AI· 8 Jul

Summary We're excited to introduce Provisioned Throughput, reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA.

Together AI1 Jul

Announcing our $800M Series C to accelerate the shift to open-source AI

Together AI· 1 Jul

Four years ago, my co-founders and I started Together AI because we saw generative AI as a turning point for human progress.

Together AI30 Jun

Together AI at ICML 2026: frontier research across the full stack

Together AI· 30 Jun

Nine papers is a lot to take in as a list. The better way to read them is by where they sit in the stack. Frontier AI is not built at a single layer.

Together AI23 Jun

ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

Together AI· 23 Jun

Summary LLMs have gotten surprisingly good at writing GPU kernels [1][2][3] , but almost all current benchmarks measuring that progress are single-GPU.

Together AI17 Jun

Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less

Together AI· 17 Jun

Summary We ran 12 landing pages through Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on nearly every page.

Together AI10 Jun

Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification

Together AI· 10 Jun

Summary Together AI has received an ISO 27001:2022 certification from A-LIGN Compliance and Security, Inc., an ANAB-accredited certification body, confirming that our Information Security Management…

Together AI2 Jun

Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets

Together AI· 2 Jun

Together AI is the preferred cloud partner for MiniMax M3. Together AI will host the open-weights model as a developer endpoint upon its public release.

Together AI29 May

How Together AI built the world’s fastest speech-to-text stack

Together AI· 29 May

‍ Artificial Analysis reported speed factor (Input audio seconds transcribed per second) -- Higher is better Modality matters A 1M-token text prompt can fit the entire Harry Potter series and still…

Together AI19 May

Benchmarking inference at scale: coding agents

Together AI· 19 May

Summary On a production coding agent workload, Together Inference Engine delivers 31% more TPS than the next fastest OSS engine on the same hardware, and maintains 2× better TTFT at saturation.

Together AI15 May

Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

Together AI· 15 May

Together AI’s mission is to make AI accessible and widely distributed. Today, we are excited to announce an exclusive partnership with Pearl Research Labs, the creators of the Pearl Network, a…

Together AI14 May

Violin: An open-source video translation skill that breaks language barriers

Together AI· 14 May

Video has become one of the most popular mediums for information sharing. Yet, the language distribution of popular video contents on the internet does not necessarily reflect the diversity of global…

Together AI12 May

Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices

Together AI· 12 May

Summary Voice finder : Search 600+ voices across MiniMax, Cartesia, Deepgram, Rime, and other models available through Together AI.

Together AI11 May

Serving DeepSeek-V4: why million-token context is an inference systems problem

Together AI· 11 May

Benchmark tables miss the main point of DeepSeek-V4: the important change is architectural. V4 turns million-token context into a serving-systems problem.

Together AI8 May

Deploy and inference any model from HuggingFace

Together AI· 8 May

Something real is shifting in how developers work. Agents open up work that used to be off-limits, not because it was technically impossible, but because it required niche expertise most of us didn't…

Together AI4 May

Foundational research powering efficient inference at scale

Together AI· 4 May

AI has spent years in the spotlight for training: the massive, GPU-intensive process of building models.

Together AI30 Apr

Announcing Together AI and Adaption Partnership

Together AI· 30 Apr

We’re excited to partner with Adaption to make Together Fine-Tuning available in Adaptive Data . Adaption is co-founded by Sara Hooker and Sudip Roy, both former leaders at Cohere and Google DeepMind…

Together AI30 Apr

From 732 bytes to nowhere: shutting down Copy Fail in production

Together AI· 30 Apr

Summary We were able to get ahead of Copy Fail (CVE‑2026‑31431) by treating it as a fleet‑level emergency, shutting off the vulnerable crypto socket interface across our infrastructure within hours…

Together AI29 Apr

DeepSeek-V4 Pro now available on Together AI

Together AI· 29 Apr

What's New DeepSeek V4 Pro on Together AI: DeepSeek V4 Pro is now available on Together AI with a 512K-token context window for long-context reasoning workloads.

Together AI28 Apr

Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0

Together AI· 28 Apr

NVIDIA Nemotron™ 3 Nano Omni is now available on the Together AI platform. Representing a meaningful step forward for multimodal AI, Nemotron 3 Nano Omni is a single, open model that reasons across…

Together AI24 Apr

Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding

Together AI· 24 Apr

Summary Distribution-aware speculative decoding (DAS) is a novel framework that significantly alleviates the rollout bottleneck in RL post-training — delivering up to 50% speedup without touching…

Together AI21 Apr

Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams

Together AI· 21 Apr

Summary Multi-tenant GPU clusters let AI-native companies share compute capacity across teams without sacrificing isolation or control.

Together AI15 Apr

Parcae: Doing more with fewer parameters using stable looped models

Together AI· 15 Apr

Summary We present Parcae, one of the first stable architectures for looped language models , achieving the quality of a Transformer twice the size with clean, predictable training.

Together AI13 Apr

EinsteinArena: Harnessing the collective intelligence of agents in the wild to advance science

Together AI· 13 Apr

Summary Scientific discovery has driven human progress, but tackling today’s hardest problems requires collective intelligence beyond any single researcher or model.

Together AI7 Apr

What is an AI Native Cloud?

Together AI· 7 Apr

Over the last few years of powering and partnering with the fastest scaling AI-native companies, we have come to realize they need a different kind of cloud: an AI Native Cloud.

Together AI3 Apr

AI for Systems: Using LLMs to Optimize Database Query Execution

Together AI· 3 Apr

Summary We worked in collaboration with Stanford University, the University of Wisconsin–Madison, and Bauplan to test whether LLMs can optimize database query execution plans.

Together AI3 Apr

Wan 2.7 video model suite now available on Together AI

Together AI· 3 Apr

Summary Four-model suite: Wan 2.7 brings video generation, continuation, and editing to Together AI, starting today with text-to-video and expanding soon to image-to-video, reference-to-video, and…

Together AI2 Apr

Deepgram speech-to-text and voice models now available natively on Together AI

Together AI· 2 Apr

Deepgram Nova-3 , Nova-3 Multilingual , Flux , and Aura-2 now run natively on Together AI Dedicated Model Inference Deepgram covers both ends of the voice pipeline, from transcription to synthesis,…

Together AI1 Apr

Inside the Together AI kernels team

Together AI· 1 Apr

The breakthrough came on a holiday weekend. Memorial Day 2022. While most of Silicon Valley was at barbecues, Dan Fu, Tri Dao, and their colleagues were about to prove the AI establishment wrong.

Together AI31 Mar

Aurora

Together AI· 31 Mar

Summary Speculative decoding goes stale in production — draft models can drift and offline retraining can't always keep pace. Aurora fixes this.

Together AI26 Mar

Plan, divide, and conquer: How weak models excel at long context tasks

Together AI· 26 Mar

TL;DR The Intuition: Don't ask one genius to read a library in an hour. Ask ten interns to read one book each.

Together AI18 Mar

Together AI expands fine-tuning service with tool calling, reasoning, and vision support

Together AI· 18 Mar

What’s New Tool call fine-tuning: Ensure agents execute structured actions reliably with end-to-end fine-tuning and inference on OpenAI-compatible schema.

Together AI17 Mar

Mamba-3

Together AI· 17 Mar

tl;dr Mamba-3 is a new state space model (SSM) designed with inference efficiency as the primary goal — a departure from Mamba-2, which optimized for training speed.

Together AI16 Mar

Together AI at NVIDIA GTC 2026: Explore our latest innovations across research and products

Together AI· 16 Mar

This year, Together AI is excited to be part of NVIDIA GTC with multiple major announcements and conversations shaping the AI ecosystem — from cutting-edge model releases to new voice AI…

Together AI12 Mar

Build real-time voice agents on Together AI

Together AI· 12 Mar

Summary Together AI, the AI Native Cloud , announced a full suite of capabilities for building real-time voice agents — co-located STT, LLM, and TTS on one cloud, eliminating inter-vendor network…