Public voice AI benchmarks increasingly show models reaching human-level performance. However, these scores do not always reflect how models behave in the real world. Because public benchmarks are open and widely used, models may gradually optimize for the tests themselves. Their improved scores may result from learning benchmark-specific patterns rather than from genuine improvements in their underlying capabilities.
One reason is that traditional benchmarks overlook many of the conditions and characteristics that determine whether a speech system is reliable, natural, contextually appropriate, and effective in practice. For this reason, we recently introduced held-out sets in Real World VoiceEQ, the Open-ASR Leaderboard, and the Far-field ASR Leaderboard to measure more of the factors that genuinely matter in real-world applications.
However, simply expanding the scope of measurement does not solve the problem. Sometimes referred to as benchmark optimization or “benchmaxxing,” this phenomenon is frequently discussed in machine learning but has remained difficult to quantify in speech recognition.
Our latest research introduces three tests to help quantify this phenomenon. We evaluated 11 widely used open-source automatic speech recognition (ASR) models and found that some of the highest-scoring systems reproduced benchmark reference transcriptions from the English portions of VoxPopuli and LibriSpeech (clean and other)—even when the audio contradicted the transcription, relevant words had been muted, or the audio equally supported two different written forms.
In some cases, models appear to rely not only on the speech content but also on subtle acoustic cues to determine which benchmark they are being evaluated on. As a result, their scores overstate their capabilities on broader speech transcription tasks.
Reference Transcription Inconsistencies (VoxPopuli Case Study)
VoxPopuli is known to contain a large number of transcription errors, which is why Artificial Analysis released a cleaned version. Our consensus discrepancy test examines how leading ASR models behave when they encounter these errors: Do they accurately transcribe what is said in the audio, or do they reproduce the erroneous reference transcription used by the benchmark?
To conduct the test at scale, we used an ensemble of multiple independent models selected based on lower phoneme error rates (PER). PER measures the degree of correspondence between a written transcription and the speech in the audio, making it a useful proxy for whether a model faithfully transcribes what it hears. The ensemble results helped us flag cases in which models consistently differed from the benchmark reference transcription. We then compared a subset of the flagged cases with human annotations to validate the corrected transcriptions.
For example, one VoxPopuli audio clip clearly contains the phrase “Thank you, Mr. President,” but the reference transcription omits “Thank you.” Of the 11 models we tested, 6 reproduced the benchmark’s erroneous transcription—providing the “expected” answer even though it contradicted the audio. In this real recording, formatting followed the same pattern as well: models that omitted “Thank you” also reproduced the benchmark’s punctuation style, writing “Mr” without a period, while models that included the audible phrase generally wrote “Mr.” with a period.
This behavior often weakened or disappeared when we used newly collected voices from recordings of members of the European Parliament or had a generic voice deliver the same content. In the example below, all but one model reverted to faithfully transcribing the audio when given a cloned version of a new parliamentary recording. This suggests that models may be responding to acoustic cues that help identify which benchmark they are being evaluated on, producing the expected transcription even when it contradicts the audio.
The reference transcription for this audio is: “Mr President, I have another complaint about this procedure, which is that it is not secret.” However, the three audio clips below actually say the same thing, with a clearly audible “Thank you,” added at the beginning of each. These cloned clips are text-to-speech (TTS) renditions of the real utterance, so the greeting can be heard in all three clips. Green highlighting and ✅ indicate that the transcription includes the audible “Thank you”; red highlighting and ❌ indicate that the transcription reproduces the benchmark’s erroneous omission of the phrase. All transcriptions are the models’ raw outputs, without any normalization—capitalization and punctuation are preserved exactly as generated by each model, including lowercase text produced by some models.
Original VoxPopuli recording
Voice clone of the same speaker
Clone of a parliamentarian speaking, recorded after the training cutoff dates of all models
| Model | Real audio | Clone of the same speaker | ep-fresh clone audio |
|---|---|---|---|
| CohereLabs/cohere-transcribe-03-2026 | ❌ Mr President… | ❌ Mr President… | ✅ Thank you, Mr President… |
| nvidia/canary-qwen-2.5b | ❌ Mr President… | ❌ Mr President… | ✅ Thank you Mr. President… |
| ibm-granite/granite-speech-4.1-2b | ❌ mr president… | ❌ mr president… | ✅ thank you mr president… |
| microsoft/Phi-4-multimodal-instruct | ❌ Mr President… | ❌ Mr President… | ❌ Mr President… |
| nvidia/parakeet-tdt-0.6b-v2 | ❌ Mr President… | ✅ Thank you, Mr President… | ✅ Thank you, Mr. President… |
| bosonai/higgs-audio-v3-8b-stt-v2 | ❌ mr president… | ❌ mr president… | ✅ thank you mr president… |
| Qwen/Qwen3-ASR-0.6B-hf | ✅ Thank you, Mr. President… | ✅ Thank you, Mister President… | ✅ Thank you, Mister President… |
| mistralai/Voxtral-Mini-3B-2507 | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… |
| moonshotai/Kimi-Audio-7B-Instruct | ✅ Thank you, mr. President… | ✅ Thank you, Mr. President… | ✅ Thank you, mr. President… |
| openai/whisper-large-v3 | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… |
| moonshine-ai/moonshine-streaming-medium | ✅ thank you mr president… | ✅ thank you mr president… | ✅ thank you mr president… |
| Number of the 11 models omitting the greeting (❌) | 6 | 5 | 1 |
Parakeet was the only model that reproduced the benchmark result on the real audio but transcribed the cloned audio of the same speaker correctly. Phi-4 was the only model that continued to omit the greeting in the ep-fresh cloned audio. When we instead resynthesized the sentence using a generic TTS voice unrelated to any parliamentary recording, all 11 models…