Speech processing
Speech-to-Text Benchmark Checklist: Accuracy, Latency, and Language Coverage
A reproducible checklist for comparing speech recognition systems with word error rate, language slices, streaming latency, and real-world audio.
Speech recognition benchmarks often compress a messy user experience into one word error rate. That number is useful, but it cannot tell you whether a model handles your accents, noisy recordings, names, or latency target.
This article explains how to compare systems. It does not claim new benchmark results or name a universal winner.
Build a test set that resembles production
Collect short and long clips across the microphones, rooms, background noise, accents, and languages you expect. Include speech with numbers, names, specialist vocabulary, silence, and overlapping voices if those matter. Separate scripted and spontaneous speech. Mozilla Common Voice offers public speech datasets that can supplement a test set, but a public corpus alone cannot represent your own users.
Keep training and prompt tuning audio away from the final test set. Record the dataset version, clip count, duration, sampling format, language labels, and permission to use each recording.
Measure transcription error consistently
Word error rate (WER) counts substitutions, deletions, and insertions divided by the number of reference words. Lower is better. Hugging Face Evaluate’s WER implementation documents the formula. Agree on normalization before scoring: case, punctuation, number formatting, and spelling variants can otherwise change the result without changing what a reader understands.
Report WER by language and recording condition, and include the underlying clip counts. For languages without clear whitespace word boundaries, add an appropriate character or token measure and explain how text was segmented. Review a sample of errors manually; one missed medication name can matter more than several harmless punctuation differences.
Measure speed as users experience it
For batch transcription, record total processing time and real time factor: processing seconds divided by audio seconds. For live captioning, measure time to first partial result and time to a stable final result. Run the same clips on the intended hardware or service tier, with identical concurrency and model settings. Warm up each system and repeat runs so one cold start does not decide the result.
Check the whole workflow
Language identification, voice activity detection, diarization, and punctuation may sit outside the recognizer. If your product needs them, test the full pipeline and report where each error begins. A detector that drops the first word will make a strong recognizer look weak. A diarization error may make a correct transcript unusable in a meeting summary.
Minimum report to publish
- Model and service versions, settings, hardware, and date
- Dataset version, language and condition counts, and normalization rules
- WER or a justified alternative for every important slice
- Batch speed or streaming latency under a stated load
- Examples of high impact failures and confidence intervals where appropriate
Use the results to choose for your actual workload, then rerun the test when a model or API version changes. See our speech-to-text comparison for model results.