Briefing

Case Studies Partner Tuning the harness, not the model: a Nemotron 3 Ultra playbook

ai-dev

Tune your agent harness by iterating on evals, focusing on system prompt, tool descriptions, and middleware, and use LangSmith traces to guide changes.

What to do now

Apply harness tuning by running a cheap eval screen, making targeted changes, and promoting only those that consistently improve scores across trials.

Summary

In a case study on harness tuning, Nemotron 3 Ultra achieved a Deep Agents suite score of 0.86—nearly matching Opus 4.8’s 0.87—while costing only $4.48 per run, a 10× cost advantage over the full suite’s $43.48. The improvement came entirely from tuning the harness—system prompt, tool descriptions, and middleware—without altering the model weights. Harness tuning follows a data‑driven eval loop: evaluate, observe, diagnose, engineer, and re‑evaluate, ensuring changes repeat across trials and do not regress other tasks. The process starts with a cheap, representative screen before running the full suite to keep costs low.

Short, single‑purpose blocks in the system prompt and context‑engineered middleware proved more effective than broad rewrites, and middleware can enforce limits, retries, and inject context at the right moments. The study shows that harness tuning can unlock frontier‑level performance from open models, but it also has a ceiling: it fixes scaffolding failures but cannot add new capabilities beyond the model’s weights. By iterating on evals and using LangSmith traces to guide changes, teams can build agents that spend their capability on the task rather than fighting the scaffolding. The case study demonstrates that a matched harness lets the model focus on the work, improving quality, cost, latency, and governance.

Key changes

  • Nemotron 3 Ultra reached a Deep Agents suite score of 0.86 with a harness tuned for the model, matching Opus 4.8’s 0.87.
  • The tuned harness achieved the score at $4.48 per run, ~10× lower cost than the $43.48 of the full suite.
  • Harness tuning improves performance without changing model weights, focusing on system prompt, tool descriptions, and middleware.
  • The eval loop involves evaluate, observe, diagnose, engineer, and re‑evaluate, with changes validated across trials.
  • A cheap, representative screen is used before the full suite to keep testing costs low.
  • Short, single‑purpose blocks in the system prompt and context‑engineered middleware are more effective than broad rewrites.
  • Middleware can enforce limits, retries, and inject context at the right moments to guide the model.
  • Harness tuning has a ceiling: it fixes scaffolding failures but cannot add new capabilities beyond the model’s weights.

Affects

enterprise

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting