arXiv:2601.16023eess.AScs.HC2026-01

用大模型实现多语言语音直译,还能保留说话人音色。

Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs

  • 用大语言模型直接处理语音翻译,跳过传统多步流程。
  • 在1000小时合成数据上训练,支持法语到英语等多语言对。
  • 加入音色控制合成技术,说话人特征和自然度显著提升。

直接语音到语音翻译(S2ST)因其减少误差传播与延迟而受到关注,但现有系统仍面临语义-声学对齐不稳、说话人身份保留困难及多语言扩展性差等问题。本文提出DS2ST-LM,一个基于多语言大语言模型(LLM)的单阶段可扩展框架,集成Whisper语音编码器、可学习投影模块、Qwen2-0.5B LLM与音色可控声码器。构建了1000小时双语语料GigaS2S-1000,通过高质量合成目标语音缓解数据稀缺问题。研究对比了语音导出的S3标记与大模型生成的文本标记两种策略,分析其对训练稳定性和语义一致性的影响。评估三种投影结构(Linear、Conv1D-Linear、Q-Former),发现简单线性投影虽收敛慢,但性能更优。实验表明,DS2ST-LM在词汇(BLEU、METEOR)与语义(BLEURT、COMET)指标上均优于传统级联与ST+TTS基线,并支持法语、西班牙语、德语、印地语、孟加拉语、乌尔都语等多语言对。进一步引入音色感知合成,使系统在说话人相似度与听感自然度上超越现有直接S2ST方案。

原文摘要 · Abstract (English)

Direct Speech-to-Speech Translation (S2ST) has gained increasing attention for its ability to translate speech from one language to another, while reducing error propagation and latency inherent in traditional cascaded pipelines. However, existing direct S2ST systems continue to face notable challenges, including instability in semantic-acoustic alignment when parallel speech data is scarce, difficulty in preserving speaker identity, and limited multilingual scalability. In this work, we introduce DS2ST-LM, a scalable, single-stage direct S2ST framework leveraging a multilingual Large Language Model (LLM). The architecture integrates a Whisper speech encoder, a learnable projection module, a Qwen2-0.5B LLM, and a timbre-controlled vocoder. We construct GigaS2S-1000, a 1000-hour bilingual corpus by extending the GigaST dataset with high-fidelity synthetic target speech, and show that this synthetic data alleviates data scarcity to some extent. We investigate two semantic token generation strategies: speech-derived S3 tokens and text-derived tokens generated by a pre-trained LLM, and analyze their impact on training stability and semantic consistency. We further evaluate three projection architectures (Linear, Conv1D-Linear, and Q-Former) and observe that while higher-capacity projectors converge faster, the simple Linear projector achieves higher performance. Extensive experiments demonstrate that DS2ST-LM outperforms traditional cascaded and ST (Qwen-Audio) + TTS baselines across both lexical (BLEU, METEOR) and semantic (BLEURT, COMET) metrics, while extending to multiple language pairs, including French, Spanish, German, Hindi, Bengali, and Urdu. Furthermore, we incorporate timbre-aware speech synthesis to preserve speaker information, enabling DS2ST-LM to surpass prior direct S2ST systems in both speaker similarity and perceptual naturalness.

语音翻译大模型音色保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。