arXiv:2608.02138cs.CL2026-08

发现语音翻译中不流畅表达含重要语义,清理反而降低质量

The Role of Disfluencies in Speech Translation

论文配图:The Role of Disfluencies in Speech Translation
图 1 · 摘自论文原文
  • 构建含标注的多语言不流畅语音翻译基准Uh-Mazing
  • 假起步和自我修正导致翻译损失,清理会直接丢弃而非误译
  • 推理时解码可缓解问题,无需重新训练模型

当前语音翻译系统(包括SpeechLLMs)在清洗后的文本上训练,倾向于去除填充停顿、假起步等不流畅表达,但此举付出代价:这些不流畅成分携带语义信息。为系统研究该问题,我们提出Uh-Mazing,一个涵盖英语至八种目标语言的人工翻译、不流畅标注的Switchboard语音数据集。在多种架构下,我们发现假起步与自我修正比填充停顿或话语标记更显著影响翻译质量;且无法保留不流畅性的模型更倾向直接省略,而非错误翻译。我们证明推理阶段解码策略可有效缓解此问题,且已公开基准数据与代码。

原文摘要 · Abstract (English)

Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them. We show this comes at a cost: disfluencies carry meaning that gets lost when speech is cleaned up. To study this systematically, we introduce Uh-Mazing, a benchmark of human-translated, disfluency-annotated Switchboard speech covering English into eight target languages. Across these languages and several architectures, we find that false starts and self-repairs, not filled pauses or discourse markers, drive most of the translation-quality loss, and that models which fail to preserve a disfluency tend to omit it rather than mistranslate it. We show inference-time decoding can mitigate this without retraining, and release the benchmark and code.

语音翻译不流畅性多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。