arXiv:2510.03115cs.CLcs.SD2025-10

CoT提示在语音翻译中主要依赖文本,几乎不利用语音信息。

Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation

  • 用归因方法分析发现CoT主要依赖转录文本
  • 注入噪声文本或直接S2TT数据可提升语音利用率
  • 现有架构需显式融合声学信息

基于ASR与T2TT模块的语音到文本翻译系统存在错误传播和无法利用语调等声学线索的问题。链式思维(CoT)提示被提出以联合利用语音与转录文本,克服上述缺陷。通过归因分析、带噪声转录本的鲁棒性测试及语调感知评估发现,CoT行为仍近似级联结构,主要依赖转录文本而极少利用语音信号。简单训练干预如引入直接S2TT数据或注入噪声转录本,能增强系统鲁棒性并提高对语音的注意力。研究结果质疑了CoT的预期优势,强调需要显式整合声学信息的新型架构。

原文摘要 · Abstract (English)

Speech-to-Text Translation (S2TT) systems built from Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) modules face two major limitations: error propagation and the inability to exploit prosodic or other acoustic cues. Chain-of-Thought (CoT) prompting has recently been introduced, with the expectation that jointly accessing speech and transcription will overcome these issues. Analyzing CoT through attribution methods, robustness evaluations with corrupted transcripts, and prosody-awareness, we find that it largely mirrors cascaded behavior, relying mainly on transcripts while barely leveraging speech. Simple training interventions, such as adding Direct S2TT data or noisy transcript injection, enhance robustness and increase speech attribution. These findings challenge the assumed advantages of CoT and highlight the need for architectures that explicitly integrate acoustic information into translation.

语音翻译链式思维声学融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。