Briefing

AI Evals Are Becoming the New Compute Bottleneck

ai-dev

Evaluation costs for agent benchmarks can reach tens of thousands of dollars; use coarse‑to‑fine strategies like Flash‑HELM or anchor‑point subsampling to cut compute by 100×–200× while preserving ranking.

What to do now

Implement coarse‑to‑fine evaluation, use Flash‑HELM or anchor‑point subsampling, and monitor cost per run to keep budgets under control.

Summary

The cost of AI evaluation has surged, with the Holistic Agent Leaderboard (HAL) spending $40,000 on 21,730 rollouts across nine models and benchmarks as of April 2026. A single GAIA run on a frontier model can cost $2,829, and the HELM benchmark’s aggregate cost across 30 models reached roughly $100,000. Subsampling strategies can reduce compute by 100×–200× while preserving ranking in static benchmarks, a technique embodied by Flash‑HELM’s coarse‑to‑fine approach. However, agent benchmarks exhibit a four‑order‑of‑magnitude cost spread, and training‑in‑the‑loop benchmarks consume 3,840 H100‑hours for a full sweep. Static benchmarks can be compressed to 2% error using anchor points, but agent evaluations remain noisy and scaffold‑sensitive, making compression less effective. The article highlights that evaluation costs can exceed training costs in scientific ML, with the Well protocol alone requiring 960 H100‑hours for a single architecture. The overall message is that evaluation compute is becoming a primary bottleneck, necessitating smarter, cost‑aware evaluation pipelines.

Key takeaways include the high cost of agent rollouts, the effectiveness of coarse‑to‑fine evaluation, and the need to monitor and manage compute budgets across different benchmark types.

Key changes

  • HAL spent $40k for 21,730 rollouts across 9 models
  • Single GAIA run costs $2,829
  • HELM aggregate cost ~ $100k across 30 models
  • 100×–200× compute reduction preserves ranking in static benchmarks
  • Flash‑HELM implements coarse‑to‑fine evaluation
  • Agent evals cost spread spans four orders of magnitude
  • Training‑in‑the‑loop benchmarks consume 3,840 H100‑hours
  • Static benchmarks can be compressed to 2% error with anchor points

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting