arXiv:2509.21144cs.SDcs.AI2025-09被引 5

UniSS实现语音到语音翻译,保留原声情感与语调。

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

  • 单阶段框架融合语音语义与风格建模,对接文本大模型。
  • 在44.8小时数据集上表现超越旧方法,保持语音一致性。
  • 适合语音翻译、语音克隆与多语言交互系统研究者。

表达性语音到语音翻译(S2ST)的终极目标是准确翻译口语内容,同时保留说话人身份与情感风格。然而,该领域进展受限于三大挑战:带有情感风格的成对语音数据稀缺、多阶段处理流程复杂,以及大语言模型(LLM)翻译能力难以迁移至语音。本文提出UniSS,一种全新的单阶段表达性S2ST框架。通过精心设计的语音语义与风格建模,实现与现有文本级大模型框架的无缝集成,构建统一的文-语音语言模型。为将文本翻译能力迁移到语音,我们提出跨模态思维链提示机制,逐步对齐音频语义与文本,并确保解码结果中的风格一致。此外,我们构建并发布了大规模高质量表达性S2ST数据集UniST,包含44.8万小时数据。实验表明,UniSS在翻译保真度和语音质量方面显著优于以往方法,同时保持语音、情绪与时长的一致性。本工作建立了下一代表达性S2ST系统的更简、更高效范式。音频样例可访问 https://cmots.github.io/uniss-demo。

原文摘要 · Abstract (English)

The ultimate goal of expressive speech-to-speech translation (S2ST) is to accurately translate spoken content while preserving the speaker identity and emotional style. However, progress in this field is largely hindered by three key challenges: the scarcity of paired speech data that retains expressive styles, the complexity of multi-stage processing pipelines, and the limited transfer of translation capabilities from large language models (LLMs). In this work, we address these challenges by introducing UniSS, a novel single-stage framework for expressive S2ST. Our approach features carefully designed speech semantic and style modeling, enabling seamless integration with existing text-based LLM frameworks to develop a unified text-speech language model. To transfer translation capabilities from text to speech, we propose a cross-modal chain-of-thought prompting process that progressively aligns audio semantics with text and ensures style preservation in the decoded results. Furthermore, we construct and release a large-scale, high-quality expressive S2ST dataset, UniST, comprising 44.8k hours of data. Experimental results show that UniSS significantly outperforms previous methods in translation fidelity and speech quality while preserving voice, emotion, and duration consistency. Our work establishes a simpler and more effective paradigm for building the next generation of expressive S2ST systems. Audio samples are available at https://cmots.github.io/uniss-demo.

语音翻译风格保留大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。