Compare transcription systems on the same representative audio, with a human reference and metrics that reflect the actual downstream job.