arXiv:2609.03992cs.CLeess.AS2026-09

无需对齐的语音配音与双工对话生成框架,支持高质量长时对话合成。

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

论文配图:Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
图 1 · 摘自论文原文
  • 采用潜空间扩散模型与文本-音频跨注意力机制,实现端到端无对齐语音生成。
  • 30亿参数模型在48万小时语料上预训练,支持单次生成1分钟内语音或任意长度生成。
  • 可自然模拟对话中的换言、回应和情绪变化,显著提升语音自然度与情感一致性。

我们提出Alignment-Free Text-Audiobox(Text-AB),一种统一的高保真语音配音与双工对话合成框架。基于以流匹配为目标训练的扩散变换器,Text-AB 在三个维度上超越了Audiobox系统:首先,采用基于DAC-VAE的潜空间扩散框架,将48 kHz波形压缩为25 Hz潜序列,压缩率超过此前EnCodec表示的10倍,同时提升重建质量;其次,完全无对齐:通过现成文本编码器直接输入原始文本,利用交叉注意力学习文本-语音对齐,无需强制对齐与显式时长预测;第三,大规模扩展模型与数据:在48万小时单语语音上预训练30亿参数模型,并在跨语言配音、双工对话合成及情绪化双工对话合成三项任务上进行微调。推理阶段支持单次生成约1分钟语音,或通过多扩散方案实现任意长度生成,结合多阶段重排序策略,提升生成质量。在真实世界配音基准测试中,相比最新内部系统,显著提升韵律相似性、语音相似性、自然度与传播性。在双工对话合成中,短对话接近真人录音水平,长对话的人类相似性与表现力大幅超越当前内部模型,且原生建模换言、应答与情绪动态。情绪条件输入显著提升情绪一致性与交互质量。

原文摘要 · Abstract (English)

We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we scale model and data substantially, pretraining a 3B-parameter model on 480k hours of monolingual speech, followed by supervised fine-tuning on three downstream tasks: cross-lingual voice dubbing, full-duplex dialogue synthesis, and emotional full-duplex dialogue synthesis. At inference, Text-AB supports one-shot generation for up to ~1 min of speech and arbitrarily long-form generation via a multi-diffusion scheme, plus a multi-stage reranking strategy that enhances quality based on automated metrics. On a real-world dubbing benchmark, Text-AB delivers a step-change improvement over the latest internal dubbing system, with large gains in prosody similarity, voice similarity, naturalness, and shareability. For full-duplex dialogue synthesis, it approaches human recordings on short-form conversations and substantially outperforms the latest internal model on long-form human-likeness and expressivity, while natively modeling turn-taking, back-channeling, and emotional dynamics. For emotional dialogue synthesis, emotion conditioning significantly improves emotion alignment and emotional interaction quality over the unconditioned baseline.

语音合成扩散模型双工对话无对齐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。