arXiv:2509.14023cs.CLcs.HC2025-09中稿 · WMT2025

用语音评估机器翻译质量,更贴近真实使用场景。

Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality

  • 通过众包收集语音判断,对比文本与语音两种评估方式。
  • 语音评估结果与文本评估基本一致,部分系统差异显著。
  • 适合关注真实语音翻译体验的研究者与产品开发者。

机器翻译已取得显著进展,尤其在语音翻译和多模态方法方面。然而,翻译质量评估仍以文本为主,依赖人工阅读对比。由于许多实际应用场景(如 Google Translate 语音模式、讯飞翻译)中翻译以语音形式呈现,通过语音而非纯文本评估更具自然性。本研究对比了10个来自WMT通用机器翻译共享任务的MT系统,在Amazon Mechanical Turk上通过众包获取的文本评估与语音评估结果。研究还进行了统计显著性检验和自重复实验,验证语音评估的可靠性和一致性。结果显示,基于语音的评估排名与文本评估总体一致,但在某些系统间识别出显著差异。这归因于语音具有更丰富、更自然的表达特性,研究建议未来将语音评估纳入机器翻译评价体系。

原文摘要 · Abstract (English)

Machine Translation (MT) has achieved remarkable performance, with growing interest in speech translation and multimodal approaches. However, despite these advancements, MT quality assessment remains largely text centric, typically relying on human experts who read and compare texts. Since many real-world MT applications (e.g Google Translate Voice Mode, iFLYTEK Translator) involve translation being spoken rather printed or read, a more natural way to assess translation quality would be through speech as opposed text-only evaluations. This study compares text-only and audio-based evaluations of 10 MT systems from the WMT General MT Shared Task, using crowd-sourced judgments collected via Amazon Mechanical Turk. We additionally, performed statistical significance testing and self-replication experiments to test reliability and consistency of audio-based approach. Crowd-sourced assessments based on audio yield rankings largely consistent with text only evaluations but, in some cases, identify significant differences between translation systems. We attribute this to speech richer, more natural modality and propose incorporating speech-based assessments into future MT evaluation frameworks.

机器翻译语音评估众包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。