不依赖参考文本,直接用语音信号评估语音识别结果。
Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

- 用预训练语音合成模型计算文本对语音的条件概率,量化声学差异。
- 在噪声环境下可降低20%错误率,显著提升识别效果。
- 无需额外训练,适合用于语音识别后处理与质量评估。
自动语音识别系统通常依赖参考转录文本进行评估,而无参考方法多依赖内部置信度或辅助语言模型。本文提出READ(基于声学差异的无参考假设评估),一种直接从语音信号评估语音识别结果的新指标。READ强调假设的声学一致性,利用预训练自回归语音合成模型,计算给定文本假设下语音标记的条件似然,以衡量语音与文本间的细粒度声学差异。无需额外训练,READ可用于假设优化。实验表明,READ与特定识别错误相关,并能改进语音识别输出,在噪声条件下相对错误率最高降低20%。
原文摘要 · Abstract (English)
Automatic speech recognition systems commonly rely on reference transcriptions for evaluation, while reference-free approaches often depend on internal confidence estimation or auxiliary language models. We propose READ (Reference-free Hypothesis Evaluation with Acoustic Discrepancy), a novel metric that evaluates ASR hypotheses directly from the speech signal. READ emphasizes the acoustic grounding of hypotheses. It uses a pretrained auto-regressive TTS model to compute the conditional likelihood of speech tokens given a text hypothesis, to measure fine-grained acoustic discrepancy between speech and text. Without additional training, READ can be applied for hypothesis refinement. Experiments show that READ correlates with specific recognition errors and improves ASR outputs, achieving up to 20\% relative error rate reduction, with particularly strong gains under noisy conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。