Briefing

LangSmith Observability & Evals: Fix Your Coding Agent Bill Doubling

ai-dev
Claude

Patch your teams to integrate LangSmith observability, enable Engine, and configure LLM Gateway cost caps.

What to do now

Patch your teams to integrate LangSmith observability, enable Engine, and configure LLM Gateway cost caps.

Summary

LangSmith now offers a unified trace model that aggregates sessions from Claude Code, Codex, Cursor, GitHub Copilot Chat, Pi, and OpenCode into a single, consistent view. The new Engine component automatically analyzes these sessions, surfaces concrete skill‑improvement recommendations, and flags redundant tool calls that waste tokens. LLM Gateway introduces cost caps that can be applied at the user, team, or organization level and will soon support routing to open‑source models for routine work. Together, these features give teams a single dashboard to see token usage, cost per session, tool calls, and sub‑agent activity across all coding agents. The observability layer is designed for teams that use more than one agent, while the Engine and Gateway plug into the same trace data to move from visibility to governance without ripping out existing workflows. The solution is built for rapid adoption: start with observability, then add Engine and LLM Gateway as needed. By July 2026, many mid‑size startups reported a 6× bill increase due to token‑maxxing, and LangSmith’s new tooling helps prevent that.

The platform’s new features are released as part of LangSmith’s ongoing product evolution, not as a major version bump, and they are available to all users immediately. They are especially relevant for enterprises that run multiple coding agents and need fine‑grained cost control.

Key changes

  • Unified trace model for Claude Code, Codex, Cursor, Copilot Chat, Pi, and OpenCode
  • Engine surfaces concrete skill‑improvement recommendations
  • LLM Gateway introduces user/team/org‑level cost caps and future open‑source routing
  • Standardized metadata and query syntax across all agents
  • Single dashboard for token usage, cost per session, tool calls, and sub‑agent activity
  • Engine flags redundant tool calls and recommends consolidations
  • LLM Gateway supports audit trails and network‑denial policies

Affects

enterprise internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting