arXiv:2501.06394cs.SDcs.AI2025-01EMNLP被引 3

统一多模态语音生成,让合成语音更贴合描述。

Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation

  • 用统一模型整合多种语音描述模态,通过对比学习对齐声音空间。
  • 在五个任务中超越原有专用模型,提升语音适配度与多样性。
  • 适合语音合成、多模态控制研究者,可直接试听样例。

个性化语音生成近年来显著提升了合成语音的真实感,但多模态驱动的语音生成仍具挑战。本文提出UniSpeaker,一种统一的多模态驱动语音生成方法。核心是基于KV-Former的统一语音聚合器,采用软对比损失将不同语音描述模态映射至共享语音空间,使生成语音更贴合输入描述。为评估多模态语音控制效果,我们构建首个多模态语音控制(MVC)基准,聚焦语音适配度、多样性与语音质量。UniSpeaker在该基准的五个任务中均优于先前的模态专用模型。语音样本可访问 https://UniSpeaker.github.io。

原文摘要 · Abstract (English)

Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generation remains on the rise. This paper introduces UniSpeaker, a unified approach for multimodality-driven speaker generation. Specifically, we propose a unified voice aggregator based on KV-Former, applying soft contrastive loss to map diverse voice description modalities into a shared voice space, ensuring that the generated voice aligns more closely with the input descriptions. To evaluate multimodality-driven voice control, we build the first multimodality-based voice control (MVC) benchmark, focusing on voice suitability, voice diversity, and speech quality. UniSpeaker is evaluated across five tasks using the MVC benchmark, and the experimental results demonstrate that UniSpeaker outperforms previous modality-specific models. Speech samples are available at \url{https://UniSpeaker.github.io}.

语音生成多模态统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。