Briefing

LangChain Test Run Comparisons

ai-dev

Enable Test Run Comparisons in LangSmith to compare multiple test runs side‑by‑side and drill into differences.

What to do now

Set up a dataset, run tests, and use the Compare view to analyze differences.

Summary

LangChain has released Test Run Comparisons, a new feature in LangSmith that lets developers compare multiple test runs side‑by‑side. The feature was built after observing that users wanted to see how new iterations performed against previous ones. It adds a UI where you can select two or more test runs from a dataset and click Compare. The view displays inputs, reference outputs, actual outputs, eval metrics, time, and latency for each datapoint.

A sidebar drill‑down with carets lets you flip between runs, and column filters let you isolate correct or incorrect cases. The comparison view supports both LLM‑assisted and regex or other evals, giving developers a quick way to spot regressions or improvements. The feature is currently in private beta but will roll out to more users in the coming weeks. It aims to improve debugging, understanding, and iteration speed for LLM applications.

Key changes

  • Added UI to select multiple test runs from a dataset and click Compare.
  • Comparison view shows inputs, reference outputs, actual outputs, eval metrics, time, and latency.
  • Sidebar drill‑down with carets allows flipping between runs.
  • Column filters let you isolate correct or incorrect cases.
  • Supports LLM‑assisted evals as well as regex or other evals.
  • Feature is in private beta with planned wider rollout.
  • Enables quick spotting of regressions or improvements.
  • Improves debugging, understanding, and iteration speed for LLM apps.

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting