提升英译中语音翻译中的重音传递能力,解决关键语义线索丢失问题。
Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation

- 构建中文重音标注数据集与基于XLS-R的重音检测器
- 提出新客观评估指标,与人工判断高度相关
- 改进CosyVoice3模型,显著提升重音保留效果
语音到语音翻译(S2ST)系统在语义准确性和语音自然度方面已取得显著进展,但跨语言重音传递——这一体现强调和说话意图的重要线索——仍严重缺乏研究,且缺乏对汉语等声调语言的可靠自动评估指标。本文通过构建带重音标注的中文数据集和基于XLS-R的普通话重音检测器,结合英语EmphAssess系统,提出一种新型跨语言重音评估客观指标。此外,我们微调CosyVoice3以构建重音感知型S2ST系统。实验表明,所提架构在重音传递能力上显著优于现有系统,同时保持了良好的翻译质量。该评估指标与人工主观评价具有强相关性。
原文摘要 · Abstract (English)
Speech-to-speech translation (S2ST) systems have achieved impressive progress in semantic accuracy and speech naturalness. However, the cross-lingual transfer of lexical stress, a vital cue for emphasis and speaker intent, remains heavily underexplored, compounded by a lack of reliable automatic evaluation metrics for tonal languages like Chinese. We investigate English-to-Chinese S2ST stress transfer by constructing a stress-annotated Chinese dataset and an XLS-R-based Mandarin stress detector. Integrating this with the English EmphAssess system, we propose a novel objective metric for cross-lingual stress evaluation. Furthermore, we fine-tune CosyVoice3 to build a stress-aware S2ST system. Experiments demonstrate that our proposed S2ST architecture significantly outperforms existing systems in stress translation capability while maintaining competitive translation quality. Furthermore, our evaluation metric exhibits a strong correlation with human subjective judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。