让语音随场景变化,实现视觉驱动的自然语音合成
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
- 构建跨模态对齐数据集Vivid-210K,打通视觉、说话人与音频关系
- 设计D-MSVA模块,实现音色与环境声学特征的精准视觉引导
- 在音质、语义清晰度和多模态一致性上超越现有模型
我们提出并定义了一项新任务——场景感知的视觉驱动语音合成,旨在解决现有语音生成模型在创造与真实物理世界一致的沉浸式听觉体验方面的局限性。针对数据稀缺与模态解耦两大挑战,我们提出VividVoice统一生成框架。首先,通过创新的程序化流水线构建了大规模高质量混合多模态数据集Vivid-210K,首次建立了视觉场景、说话人身份与音频之间的强关联。其次,设计核心对齐模块D-MSVA,利用解耦记忆库架构与跨模态混合监督策略,实现从视觉场景到音色及环境声学特征的细粒度对齐。主观与客观实验结果均表明,VividVoice在音频保真度、内容清晰度和多模态一致性方面显著优于现有基线模型。演示地址:https://chengyuann.github.io/VividVoice/
原文摘要 · Abstract (English)
We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world. To tackle the two core challenges of data scarcity and modality decoupling, we propose VividVoice, a unified generative framework. First, we constructed a large-scale, high-quality hybrid multimodal dataset, Vivid-210K, which, through an innovative programmatic pipeline, establishes a strong correlation between visual scenes, speaker identity, and audio for the first time. Second, we designed a core alignment module, D-MSVA, which leverages a decoupled memory bank architecture and a cross-modal hybrid supervision strategy to achieve fine-grained alignment from visual scenes to timbre and environmental acoustic features. Both subjective and objective experimental results provide strong evidence that VividVoice significantly outperforms existing baseline models in terms of audio fidelity, content clarity, and multimodal consistency. Our demo is available at https://chengyuann.github.io/VividVoice/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。