Judging generated audio is less about a single score than about a small set of dimensions, each of which can fail independently. The first is naturalness and prosody, whether the speech carries the rise and fall, pauses, and emphasis a human would use or flattens into a monotone that is clear but lifeless. The second is artifact-freeness. Listen for the tells: metallic ringing, clicks at word boundaries, a breath in the wrong place, a sudden shift in room tone, or a note that smears in music. These are the audio equivalent of extra fingers, and once you hear them you cannot unhear them.
The third dimension is consistency over time, and it is where long pieces break. A voice should keep the same timbre, accent, and energy from the first minute to the fortieth, and a song should hold its arrangement together rather than drifting. Short demos hide this, so always evaluate at the length you actually need. The fourth is controllability, whether you can direct emotion, pacing, and pronunciation or are stuck re-rolling until you get a usable take. The fifth is intelligibility, which sounds trivial but is not, especially for names, numbers, acronyms, and sung lyrics. The sixth is provenanceProvenanceA verifiable record of how a piece of media was made and by whom, increasingly required so that synthetic content can be traced.: increasingly, quality includes whether an output is traceable and disclosed. Google embeds SynthID watermarks by default in Lyria audio, designed to survive compression and even re-recording and detectable by software rather than by ear.
Consider how the weighting shifts. Audiobook narration lives or dies on consistency and prosody across hours and can tolerate slow, offline generation. A real-time voice agent weightsWeightsThe learned numbers inside a model that encode everything it knows. 'Open weights' means these are downloadable.latencyLatencyThe delay between asking for a generation and getting the first result. Distinct from throughput. and intelligibility above expressive range, because a warm voice that arrives late feels broken. Film dialogue demands the highest naturalness and fine emotional control, and usually has to sit convincingly next to real recordings, so timbre realism is unforgiving. A music bed under a video needs to loop and duck cleanly and stay out of the way, so structural perfection matters less than mood and mixability, which is where stemsStemsThe separated individual tracks of a piece of music, such as vocals, drums, and bass, which can be edited or remixed on their own. earn their keep. Sound effects and foleyFoleyThe everyday sound effects of a scene, footsteps, cloth, impacts, now increasingly generated from the video itself. are judged almost entirely on timing and physical plausibility: a footstep that lands a frame late reads as fake even if the sound itself is flawless. Naming the use case first, then weighting these dimensions, is the whole discipline.
Check yourself0 / 6
Q01
Why is there no single bar for audio quality?
Q02
What are some of the independent dimensions along which generated audio can fail?
Q03
Why should you always evaluate audio at the full length you actually need?
Q04
How does the weighting of dimensions shift between a real-time voice agent and film dialogue?
Q05
Why are blind preference tests and vendor demos a poor guide for real audio projects?
Q06
Why are sound effects and foley judged differently from music or narration?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.