现有语音翻译评估方法无法有效检测语音特征,论文提出新模型并揭示三大原因。
Why We Need Speech to Evaluate Speech Translation
- 用语音编码器构建SpeechCOMET,结合语音信号进行质量评估
- 新模型在标准任务上超越文本基线,但仍难捕捉语音特有信息
- 适合关注语音翻译评估与语音建模的研究者
语音翻译模型越来越能保留说话人性别、语调和强调等语音特征,但评估指标仍对此类现象视而不见。我们在两个对比数据集(聚焦性别一致性和语调)上对文本与语音质量评估指标进行了元评估,发现即使有语音信号输入,两者均表现不佳。随后我们训练了配备语音编码器的SpeechCOMET系列模型,并以先进的SpeechLLM作为评判者进行评估。两者在标准质量评估中表现匹配或超过文本基线的COMET,但在持续评估语音特定现象方面仍不理想。我们识别出三个关键原因:(1) 当前编码器无法可靠保留语音特征;(2) 模型常忽略语音源信号;(3) 质量评估训练数据中相关样本过少。所有模型与代码已开源,研究主张进步需依赖专门设计的语音特异性训练数据和真正基于语音的模型。
原文摘要 · Abstract (English)
Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet evaluation metrics remain blind to such phenomena. We meta-evaluate both text- and speech-based quality estimation metrics on two contrastive datasets targeting gender agreement and prosody, and find that both fall short, even when given direct access to the speech signal. We then train SpeechCOMET, a family of quality estimation models with speech encoders, and evaluate a state-of-the-art SpeechLLM as a judge. Both match or exceed text-based COMET on standard quality estimation, but neither consistently assesses speech-specific phenomena. We identify three causes: (1) speech-specific features are not reliably preserved in current encoders, (2) models tend to ignore the speech source signal, and (3) quality estimation training data contains too few relevant examples. We release all models and code, and argue that progress requires dedicated speech-specific training data and models that genuinely condition on speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。