清理语音中与语义无关的干扰成分,提升端到端语音翻译效果
Representation Purification for End-to-End Speech Translation
- 将语音表示分解为内容相关和无关成分,针对性去除后者
- 在MuST-C和CoVoST-2上全方向提升翻译性能,尤其在无转录文本时表现最优
- 适用于追求高鲁棒性语音翻译的场景,尤其适合资源受限的跨语言任务
语音到文本翻译(ST)是一项跨模态任务,旨在将口语转换为目标语言的文本。以往研究主要通过促进机器翻译知识迁移来提升性能,探索了多种弥合语音与文本模态差距的方法。然而,语音中与语义无关的因素(如音色、节奏)仍会阻碍知识迁移效率。本文将语音表示视为内容无关与内容相关因素的组合,通过初步实验发现,引入内容无关扰动会导致翻译性能显著下降。为此,提出带监督增强的语音表示净化框架SRPSE,通过剔除语音表示中的内容无关成分,减轻其对语音翻译的负面影响。在MuST-C和CoVoST-2数据集上的实验表明,SRPSE在三种设置下均显著提升所有翻译方向的性能,并在无转录文本(transcript-free)设置下取得最优表现。
原文摘要 · Abstract (English)
Speech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowledge transfer from machine translation, exploring various methods to bridge the gap between speech and text modalities. Despite substantial progress made, factors in speech that are not relevant to translation content, such as timbre and rhythm, often limit the efficiency of knowledge transfer. In this paper, we conceptualize speech representation as a combination of content-agnostic and content-relevant factors. We examine the impact of content-agnostic factors on translation performance through preliminary experiments and observe a significant performance deterioration when content-agnostic perturbations are introduced to speech signals. To address this issue, we propose a \textbf{S}peech \textbf{R}epresentation \textbf{P}urification with \textbf{S}upervision \textbf{E}nhancement (SRPSE) framework, which excludes the content-agnostic components within speech representations to mitigate their negative impact on ST. Experiments on MuST-C and CoVoST-2 datasets demonstrate that SRPSE significantly improves translation performance across all translation directions in three settings and achieves preeminent performance under a \textit{transcript-free} setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。