arXiv:2410.24019cs.CLcs.SD2024-10被引 12

探究语音翻译系统能否理解语调,发现端到端模型有潜力但实际效果有限。

Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?

  • 用大语言模型和可控语音合成生成对比语音样本评估语调感知能力
  • 端到端模型虽能捕捉语调但影响翻译效果不强,优于分步系统
  • 分步系统也能部分感知语调,但依赖文本表面形式且程度较弱

语音的语调特征(如重音、语调、节奏)对语义有显著影响,进而可能影响文本翻译结果。然而,语调在语音转文本翻译(S2TT)系统中极少被研究。尽管端到端(E2E)系统理论上更适合感知语调,因其直接访问语音信号做翻译决策,但其实际表现尚不明确。主要挑战在于难以评估系统对语调的感知能力。为此,本文提出一种新评估方法与聚焦式基准(ContraProST),通过大语言模型和可控文本转语音(TTS)生成对比样本。在英译德、西、日的实验中发现:(a) S2TT模型具备一定的语调内部表征,但语调信号通常不足以显著影响翻译;(b) E2E系统优于语音识别+文本翻译的级联系统,验证了其理论优势;(c) 某些级联系统也能捕捉语调信息,但程度较弱,且取决于转录文本的表面形式。

原文摘要 · Abstract (English)

The prosody of a spoken utterance, including features like stress, intonation and rhythm, can significantly affect the underlying semantics, and as a consequence can also affect its textual translation. Nevertheless, prosody is rarely studied within the context of speech-to-text translation (S2TT) systems. In particular, end-to-end (E2E) systems have been proposed as well-suited for prosody-aware translation because they have direct access to the speech signal when making translation decisions, but the understanding of whether this is successful in practice is still limited. A main challenge is the difficulty of evaluating prosody awareness in translation. To address this challenge, we introduce an evaluation methodology and a focused benchmark (named ContraProST) aimed at capturing a wide range of prosodic phenomena. Our methodology uses large language models and controllable text-to-speech (TTS) to generate contrastive examples. Through experiments in translating English speech into German, Spanish, and Japanese, we find that (a) S2TT models possess some internal representation of prosody, but the prosody signal is often not strong enough to affect the translations, (b) E2E systems outperform cascades of speech recognition and text translation systems, confirming their theoretical advantage in this regard, and (c) certain cascaded systems also capture prosodic information in the translation, but only to a lesser extent that depends on the particulars of the transcript's surface form.

语音翻译语调分析端到端模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。