A new comparison of DeepSeek V4 Pro 0813 and GPT-5.6 Sol shows that a cascading strategy—using DeepSeek V4 Pro 0813 first and escalating to GPT-5.6 Sol when tests fail—solves 83.0% of DeepSWE tasks at $3.35 each, outperforming Sol alone (72.7%) at a lower cost. The two models differ in price (35x gap), reliability, and failure modes.
From the source
Run DeepSeek V4 Pro 0813 first, escalate to GPT-5.6 Sol when the tests fail. That cascade solves 83.0% of DeepSWE tasks at $3.35 each. Sol alone solves 72.7% at $8.37. Ten points better, 60% cheaper.
together.ai