为语音翻译错误自动标注提供新方法,提升系统可信度。
Automatic Labelling of Speech Translation Errors

- 提出STEL标注协议与真实数据集,支持误差分类。
- 文本模型与多模态大模型在标注上达到人类一半精度。
- 揭示语音处理与文本模型互补性,适合评估系统缺陷。
语音翻译中的错误会降低系统可信度,可能造成严重后果。然而目前尚无统一的评估信心与质量估计的方法。为此,我们提出语音翻译错误标注(STEL)框架,建立标注规范、构建小型真实端到端评估数据集,并分析现有纯文本与语音处理系统的表现。结果表明,纯文本模型XCOMET和多模态大模型Qwen2.5-Omni在STEL任务中可达到人类约50%的精确度。研究还发现,直接处理语音对错误标注至关重要,且当前文本与语音处理系统在识别翻译错误与语音处理错误方面具有互补性。
原文摘要 · Abstract (English)
Errors in speech translations reduce trustworthiness of Speech Translation (ST) systems and can have serious consequences. Yet currently there is no established methodology for evaluating confidence and quality estimation of speech translations. To initiate progress in this direction, we propose Speech Translation Error Labelling (STEL). We create an annotation protocol, a small authentic end-to-end evaluation dataset, and we analyse how existing text-only and speech-processing systems perform the STEL task. Our results show that text-only XCOMET and multimodal LLM Qwen2.5-Omni are able to perform the STEL task in roughly half the precision of humans. We also find that direct speech processing is necessary for the STEL task, and that the current text-only and speech-processing systems are complementary in labelling translation-only vs. speech-processing errors in ST.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。