Briefing

LangSmith Agent Architecture Quickly Start Evaluating LLMs With OpenEvals

ai-dev

Integrate openevals and agentevals into your evaluation pipeline to standardize LLM quality checks.

What to do now

Integrate openevals and agentevals into your evaluation pipeline and configure LangSmith logging to capture results.

Summary

LangSmith has released two new packages—openevals and agentevals—designed to streamline the creation of LLM evaluation pipelines. The openevals package offers pre‑built LLM‑as‑a‑judge evaluators that can be customized with starter prompts and few‑shot examples, while the agentevals package focuses on agent‑specific metrics such as trajectory evaluation and tool‑call validation. Structured data evaluators are also included, allowing exact‑match checks or LLM‑as‑a‑judge validation for JSON or other structured outputs. Both packages support aggregation of scores across multiple feedback keys, providing a high‑level view of evaluator performance. Integration with LangSmith is recommended so that evaluation results are logged, shared, and visualized within the same platform used for tracing and debugging. The release emphasizes best practices for dataset curation and metric selection, and it plans to add RAG‑specific and multi‑agent evaluators in the coming weeks. Users can contribute new evaluators via GitHub pull requests, fostering a community‑driven ecosystem. The update positions LangSmith as a central hub for production‑grade LLM quality assurance.

Key changes

  • openevals provides pre‑built LLM‑as‑a‑judge evaluators with customizable prompts
  • agentevals offers agent trajectory evaluation for tool‑call sequences
  • Structured data evaluators support exact‑match and LLM‑as‑a‑judge validation
  • Integration with LangSmith logs and shares evaluation results
  • Support for evaluating hallucination, conversational quality, and writing coherence
  • Pre‑built starter prompts and few‑shot examples reduce setup time
  • Aggregation of scores across feedback keys gives high‑level performance views
  • Future plans include RAG and multi‑agent specific evaluators

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting