Briefing

Agent Execution Tax Benchmark Shows Open‑Weight Models Outperform Gemini

ai-dev
by /u/ogandrea ·

Run the Agent Execution Tax benchmark on your agents to compare cost per successful task.

What to do now

Benchmark your agents using Agent Execution Tax to identify the most cost‑effective model.

Summary

The benchmark evaluates 720 browser‑agent tasks across four models, introducing the Agent Execution Tax metric that measures wasted versus productive inference. MiniMax M2.5 is 2.3× cheaper per successful task than Gemini, while GLM‑5 achieves the highest accuracy at 57.1% and excels on structured data. Kimi K2.5 records the lowest parse retries across 852 calls, compared to Gemini’s 18.6%. The results reveal that open‑weight models win not because they are smarter but because they are more reliable per call. Token‑pricing comparisons are misleading once retries compound, underscoring the importance of the new metric.

The study also highlights that open‑weight models hold their own against Gemini 2.5 Flash, and that the Agent Execution Tax metric provides a clearer picture of cost per successful task. The benchmark is fully reproducible, with detailed steps available in the linked repository. These insights can help teams choose cost‑effective models for browser‑agent workloads.

Key changes

  • MiniMax M2.5 is 2.3× cheaper per successful task than Gemini.
  • GLM‑5 achieves 57.1% accuracy, strongest on structured data.
  • Kimi K2.5 records 0% parse retries across 852 calls, vs Gemini 18.6%.
  • Open‑weight models win due to reliability per call, not intelligence.
  • Token‑pricing comparisons mislead when retries compound.
  • 720 browser‑agent tasks benchmarked.
  • Agent Execution Tax metric measures wasted vs productive inference.
  • Open‑weight models hold against Gemini 2.5 Flash.

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting