Measurements, methods, and the cases that still fail.
We measured our models against 63 audio files and attested reference transcripts. The numbers, the outliers, and the cases that still fail.