Key Takeaways While GPT-5.6 Luna is the stronger engineer on every quality measure, DeepSeek-V4 Flash 0731 is cheap enough that a DeepSeek-first cascade beats Luna alone on both accuracy and cost. GPT-5.6 Luna leads DeepSWE pass@1 decisively at 67.2% vs 53.3%, a 14 point gap, and holds the lead at every equal attempt count. DeepSeek-V4 Flash is the cheapest model on the DeepSWE board: \$0.10 per rollout vs \$0.61, delivering 532 solves per \$100 against Luna's 110. DeepSeek-V4 Flash fails more cleanly, breaking the repo's existing test suite in 9% of failures vs Luna's 15%. Running DeepSeek-V4 Flash first and escalating to Luna only on failure solves 78.9% of tasks at \$0.385 each: more accurate than Luna alone and 37% cheaper. Available now · Open weights Run DeepSeek-V4 Flash on Together AI The cheapest model on the DeepSWE board, about ten cents a task. Serve the cheap first stage at production scale without frontier token prices. Open the playground DeepSeek-V4 Flash 0731 is the cheapest model on the entire DeepSWE board: about ten cents a task! GPT-5.6 Luna, a solid upper-tier flagship, runs \$0.61 a task, roughly six times DeepSeek’s price. …