Speech Testing
Audio and Speech Eval Frameworks: WER, BLEURT, and MOS Scoring in CI
Speech AI systems — transcription APIs, TTS engines, voice assistants — produce outputs that don't fit a binary pass/fail test. The question is never "did it produce output" but "is the output good enough." Answering that question consistently, automatically, and in CI requires evaluation frameworks