arXiv:2510.13194cs.CLcs.AI2025-10被引 2

让语音翻译保留原话重音,提升跨语言情感传递效果

StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation

  • 用大模型将源语言重音转为目标语言标签,指导可控语音合成
  • 在少量真实数据下仍显著优于基线,重音保留率更高
  • 适合需要情感和语气精准传递的语音翻译场景

我们提出一种注重重音的语音到语音翻译系统,通过大模型实现跨语言重音转换。该方法将源语言重音转化为目标语言标签,引导可控语音合成模型生成对应强调。为应对数据稀缺问题,构建了自动对齐训练数据生成管道,并引入LLM作为评估裁判。实验表明,该方法在保持相当翻译质量、说话人意图与自然度的前提下,显著提升了重音保留能力。本工作凸显了韵律在翻译中的重要性,为语音翻译中非语言信息的保留提供了高效、数据节约的解决方案。

原文摘要 · Abstract (English)

We propose a stress-aware speech-to-speech translation (S2ST) system that preserves word-level emphasis by leveraging LLMs for cross-lingual emphasis conversion. Our method translates source-language stress into target-language tags that guide a controllable TTS model. To overcome data scarcity, we developed a pipeline to automatically generate aligned training data and introduce the "LLM-as-Judge" for evaluation. Experiments show our approach substantially outperforms baselines in preserving emphasis while maintaining comparable translation quality, speaker intent, and naturalness. Our work highlights the importance of prosody in translation and provides an effective, data-efficient solution for preserving paralinguistic cues in S2ST.

语音翻译重音保留大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。