让语音翻译保留笑声与哭声,提升情感真实感
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation

- 用专家混合模型精准捕捉复合情绪表达
- 仅需30分钟数据即达优秀表现,复现率76%
- 适合注重情感传达的跨语言对话系统研发
近期语音到语音翻译(S2ST)系统虽具备高语义准确性,却普遍丢失笑声、哭泣等非语言发声(NVs),严重影响实际应用。本文提出三项贡献:第一,构建可扩展的表达性数据集合成流程,缓解数据稀缺问题;第二,提出MoVE架构,采用基于LoRA的专家混合结构,结合专用适配器与软权重路由机制,有效建模混合情感状态;第三,证明预训练音频大模型具有惊人数据效率:仅需30分钟精选数据即可达到优异性能。在英中S2ST任务上,相比强基线模型,MoVE在76%的情况下成功还原目标非语言发声,人类评估显示其自然度与情感保真度最优,而现有系统最多仅保留14%的非语言发声。
原文摘要 · Abstract (English)
Recent Speech-to-Speech Translation (S2ST) systems achieve strong semantic accuracy yet consistently strip away non-verbal vocalizations (NVs), such as laughter and crying that convey pragmatic intent, which severely limits real-world utility. We address this via three contributions. First, we propose a synthesis pipeline for building scalable expressive datasets to overcome the data scarcity limitation. Second, we propose MoVE, a Mixture-of-LoRA-Experts architecture with expressive-specialized adapters and a soft-weighting router that blends experts for capturing hybrid expressive states. Third, we show pretrained AudioLLMs enable striking data efficiency: 30 minutes of curated data is enough for strong performance. On English-Chinese S2ST, while comparing with strong baselines, MoVE reproduces target NVs in 76% of cases and achieves the highest human-rated naturalness and emotional fidelity among all compared systems, where existing S2ST systems preserve at most 14% of NVs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。