DeepSeek V4 Pro Matches GPT‑5.2 on FoodTruck Bench, 17× Cheaper
Compare DeepSeek V4 Pro to GPT‑5.2 on FoodTruck Bench: similar performance, 17× cheaper.
Explore integrating DeepSeek V4 Pro into cost‑sensitive AI workloads to reduce API spend.
Summary
DeepSeek V4 Pro, the first Chinese model to reach frontier tier on the FoodTruck Bench, matched GPT‑5.2’s median outcome within 3 % after a 10‑week gap since GPT‑5.2’s mid‑February test. The benchmark, which simulates a food‑truck operation with 34 tools and persistent memory, places DeepSeek at #4 overall behind Opus 4.6, GPT‑5.2, and Grok 4.3.
Pricing is a major differentiator: GPT‑5.2 charges $1.75/M input and $14/M output, while DeepSeek runs at $0.435/M input and $0.87/M output, making it roughly 17× cheaper for the same agentic workload. In cost‑efficiency terms, DeepSeek ranks #2 on the leaderboard, only behind Gemma 4 31B, and outperforms Grok 4.3 in consistency—zero losses, six times less food waste, 30 % more meals per day, and a 2.4× tighter outcome distribution. Opus 4.6 still achieves a higher peak, and Gemma remains the cheapest option. Xiaomi MiMo v2.5 Pro also entered the top six, posting a 1,019 % median ROI and $22,388 median net worth at $2.41/run, though its variance is wider. Together, these results show two Chinese models in the top six, both under $3.5/run, a sharp shift from the frontier gap of a year to just ten weeks.
Key changes
- DeepSeek V4 Pro matches GPT‑5.2 median outcome within 3 % on FoodTruck Bench
- DeepSeek V4 Pro is 17× cheaper: $0.435/M input vs $1.75/M input, $0.87/M output vs $14/M output
- DeepSeek ranks #2 overall on cost‑efficiency leaderboard, behind Gemma 4 31B
- DeepSeek outperforms Grok 4.3 in consistency: zero losses, six times less food waste, 30 % more meals per day, 2.4× tighter outcome distribution
- Xiaomi MiMo v2.5 Pro lands #6 with 1,019 % median ROI and $22,388 median net worth at $2.41/run