arXiv:2511.03295cs.CLcs.AI2025-11被引 3

提出用音频转录和反向翻译构建文本代理,提升语音翻译评估准确性

How to Evaluate Speech Translation with Source-Aware Neural MT Metrics

  • 用ASR转录或反向翻译生成语音的文本代理
  • 当错误率低于20%时,ASR转录更可靠,反向翻译更便宜有效
  • 新算法解决合成源与参考译文对齐问题,适合真实场景使用

语音翻译自动评估通常依赖参考译文对比,但忽略了源端信息。近年来,结合源文本的神经机器翻译评估指标在机器翻译中表现更优。将其应用于语音翻译面临挑战:源为音频,且常无可靠转录或对齐。本文首次系统研究语音翻译中的源感知评估方法,重点关注无源转录的真实场景。探索两种文本代理生成策略:语音识别转录(ASR)与参考译文的反向翻译,并提出一种两步跨语言重分割算法,解决合成源与参考译文间的对齐错位问题。在涵盖79种语言对、六种不同架构系统的两个语音翻译基准上实验表明:当词错误率低于20%时,ASR转录优于反向翻译;反向翻译则始终计算成本更低但依然有效。低资源语言对(Bemba-English)及人工质量判断验证了结果稳健性。该方法使源感知评估在语音翻译中更具鲁棒性,推动更准确、严谨的评估范式发展。

原文摘要 · Abstract (English)

Automatic evaluation of ST systems is typically performed by comparing translation hypotheses with one or more reference translations. While effective to some extent, this approach inherits the limitation of reference-based evaluation that ignores valuable information from the source input. In MT, recent progress has shown that neural metrics incorporating the source text achieve stronger correlation with human judgments. Extending this idea to ST, however, is not trivial because the source is audio rather than text, and reliable transcripts or alignments between source and references are often unavailable. In this work, we conduct the first systematic study of source-aware metrics for ST, with a particular focus on real-world operating conditions where source transcripts are not available. We explore two complementary strategies for generating textual proxies of the input audio, ASR transcripts, and back-translations of the reference translation, and introduce a novel two-step cross-lingual re-segmentation algorithm to address the alignment mismatch between synthetic sources and reference translations. Our experiments, carried out on two ST benchmarks covering 79 language pairs and six ST systems with diverse architectures and performance levels, show that ASR transcripts constitute a more reliable synthetic source than back-translations when word error rate is below 20%, while back-translations always represent a computationally cheaper but still effective alternative. The robustness of these findings is further confirmed by experiments on a low-resource language pair (Bemba-English) and by a direct validation against human quality judgments. Furthermore, our cross-lingual re-segmentation algorithm enables robust use of source-aware MT metrics in ST evaluation, paving the way toward more accurate and principled evaluation methodologies for speech translation.

语音翻译评估方法源感知ASR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。