arXiv:2604.26136eess.AScs.CL2026-04

让科学家语音跨语言克隆,保持原声辨识度。

One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

  • 基于OmniVoice模型构建跨语言语音克隆系统
  • 合成数据微调使语音可懂度与相似度提升
  • 适用于科学演讲的多语种语音迁移场景

在科学传播等专业领域,如何在改变语言的同时保留说话人语音特征,仍是语音技术中的核心挑战。本文针对国际口语翻译会议(IWSLT 2026)的跨语言语音克隆共享任务,评估了多个先进语音克隆模型在阿拉伯语、中文和法语文本上的表现。随后,基于OmniVoice基础模型构建语音克隆系统,采用来自ACL 60/60语料库的多模型集成蒸馏进行数据增强。通过实验证明,使用该合成数据进行微调可有效提升语音可懂度(WER与CER降低)及说话人相似度(SIM提升),效果在不同语言间存在差异。

原文摘要 · Abstract (English)

Preserving a speaker's voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, particularly in specialized domains such as scientific communication. In this paper, we address this challenge through our system submission to the International Conference on Spoken Language Translation (IWSLT 2026), the Cross-Lingual Voice Cloning shared task. First, we evaluate several state-of-the-art voice cloning models for cross-lingual speech generation of scientific texts in Arabic, Chinese, and French. Then, we build voice cloning systems based on the OmniVoice foundation model. We employ data augmentation via multi-model ensemble distillation from the ACL 60/60 corpus. We investigate the effect of using this synthetic data for fine-tuning, demonstrating improvements in intelligibility (WER & CER) and speaker similarity (SIM), with gains varying across languages.

语音克隆跨语言科学语音OmniVoice

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。