Briefing

FFASR Leaderboard: Benchmarking ASR in the Real World

ai-dev

Test your ASR models on the FFASR Leaderboard to quantify far‑field WER and latency trade‑offs.

What to do now

Test your ASR models on the FFASR Leaderboard and analyze the far‑field WER vs RTFx Pareto front.

Summary

On June 24 2026, Treble Technologies and Hugging Face launched the FFASR Leaderboard, the first open, community‑driven benchmark for far‑field automatic speech recognition. The leaderboard evaluates models across 14 simulated rooms that range from 20 m³ to 470 m³, covering bathrooms, living rooms, offices, classrooms, and restaurants, and validates the simulation against real‑world measurements. It tests nine acoustic conditions, with four primary ranking columns: near‑field dry, far‑field high SNR (>14 dB), far‑field mid SNR (8–12 dB), and far‑field low SNR (<6 dB), plus Lab Measured and Lab Simulated columns for sim‑to‑real validation and moving‑source splits in beta. Accuracy and latency are plotted on a Pareto front of average WER versus RTFx, evaluated on an NVIDIA L4 GPU, revealing that far‑field WER at low SNR is several times higher than near‑field WER. The leaderboard accepts Hugging Face model IDs and supports Whisper, IBM Granite, Cohere Transcribe, Wav2Vec2, HuBERT, SpeechBrain, and custom evaluators. Future tracks include multi‑talker scenarios, microphone array support, and echo cancellation. The benchmark aims to expose the gap between clean‑speech benchmarks and real‑world acoustic robustness, encouraging researchers to prioritize far‑field performance.

Key changes

  • Launch of FFASR Leaderboard (June 24 2026) by Treble Technologies & Hugging Face.
  • Benchmark covers 14 simulated rooms (20–470 m³) with 14 fully furnished rooms.
  • Evaluates nine acoustic conditions; four primary ranking columns: near‑field dry, far‑field high SNR, far‑field mid SNR, far‑field low SNR.
  • Includes Lab Measured and Lab Simulated columns for sim‑to‑real validation.
  • Accuracy vs latency Pareto front plotted using average WER vs RTFx on NVIDIA L4 GPU.
  • Current submissions reveal far‑field WER at low SNR is several times higher than near‑field WER.
  • Future roadmap: multi‑talker scenarios, microphone array support, echo cancellation.
  • Submission pipeline accepts Hugging Face model IDs and supports Whisper, IBM Granite, Cohere Transcribe, Wav2Vec2, HuBERT, SpeechBrain, and custom evaluators.

Affects

none

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting