arXiv:2606.19325cs.SDcs.AI2026-06

用自然语言描述对话场景,生成带多人声音的逼真音频。

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

论文配图:Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors
图 1 · 摘自论文原文
  • 用参考音色+自然语言提示,直接控制多说话人音频生成。
  • 在CoVoMix2上超越现有系统,支持重叠对话和环境音效。
  • 解决参考音色误匹配问题,让模型依赖文本而非声音相似性。

现有多方言对话系统依赖结构化标注(如每轮标签、多流转录)绑定说话人,仅生成纯净语音,缺乏真实对话的背景音效。本文提出ScenA方法,将一个预训练于大规模野外数据的文本到音频流匹配基础模型,直接以多个参考语音和自由形式的自然语言提示为条件,生成完整多说话人音频场景。该方法继承了基础模型对非录音棚音频的建模能力:背景噪声、房间混响、重叠对话与自发副语言事件。具体实现中,将参考隐变量拼接至模型序列,并通过轻量级身份感知位置编码区分。然而我们发现关键障碍——‘参考捷径’:标准噪声调度下,模型可通过目标语音与参考音色的声学相似性识别匹配,绕过文本提示。为此提出高噪声偏置的时间步分布,迫使模型依赖文本进行说话人分配。在CoVoMix2-Dialogue基准上评估,ScenA在说话人绑定指标上优于现有系统,同时生成包含重叠语句、情绪发声与环境音的丰富对话音频。结果表明,使用通用音频模型结合自由描述,优于传统语音流水线加结构化脚本的方式。

原文摘要 · Abstract (English)

Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations. We take a different approach. Our method, ScenA, conditions a text-to-audio flow-matching foundation model, pretrained on large-scale in-the-wild data, directly on multiple reference voices and a free-form natural language prompt that describes an entire multi-speaker audio scene. Leveraging such a foundational model allows us to inherit its capacity for natural, non-studio audio: background noise, room acoustics, overlapping dialogue, and spontaneous paralinguistic events, while adding multi-speaker control without any per-turn structure. Concretely, reference latents are concatenated into the model's token sequence and distinguished by lightweight identity-aware positional encodings. However, we identify a critical obstacle to this approach: the \textit{Reference Shortcut}. During training under standard noise schedules, the model can identify the matching reference by acoustic similarity to the noisy target, bypassing the text prompt entirely. We address this with a high-noise-biased timestep distribution that forces the model to rely on the text prompt for speaker assignment. We evaluate ScenA on the CoVoMix2-Dialogue benchmark, showing that it outperforms existing multi-speaker systems on speaker-binding metrics while generating rich conversational audio with overlapping speech, emotional vocalizations, and ambient sound. Our results demonstrate the advantage of using a general-purpose audio model conditioned on a free-form scene description, rather than passing structured dialog scripts through a speech-only pipeline.

音频生成多说话人文本驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。