arXiv:2509.24773eess.AScs.AI2025-09被引 8

统一视频驱动的音效与语音生成,提升跨模态合成效果。

VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learning

  • 用分离条件聚合机制,分治语义与时间特征输入
  • 联合训练下性能优于单任务模型,打破性能下降魔咒
  • 支持合成数据快速适配,适合多模态音频生成研究者

视频驱动的音频生成,包括视频到音效(V2S)和视觉文本到语音(VisualTTS),传统上被视为独立任务,未充分探索统一生成框架的潜力。本文提出VSSFlow,一种统一的流匹配框架,无缝解决这两类问题。为在扩散变压器(DiT)架构中有效处理多种输入信号,我们设计了一种解耦条件聚合机制,利用注意力层的不同内在特性:交叉注意力用于语义条件,自注意力用于时序密集条件。此外,与普遍认为联合训练会导致性能下降的观点相反,我们证明了VSSFlow在端到端联合学习过程中仍保持优异性能。进一步地,我们采用简单的特征级数据合成方法,表明该框架可稳健适应基于合成数据的音效与语音联合生成。在V2S、VisualTTS及联合生成基准上的大量实验表明,VSSFlow能有效统一这些任务,并超越现有领域专用基线,凸显统一生成模型的关键潜力。

原文摘要 · Abstract (English)

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In this paper, we bridge this gap with VSSFlow, a unified flow-matching framework that seamlessly solve both problems. To effectively handle multiple input signals within a Diffusion Transformer (DiT) architecture, we propose a disentangled condition aggregation mechanism leveraging distinct intrinsic properties of attention layers: cross-attention for semantic conditions, and self-attention for temporally-intensive conditions. Besides, contrary to the prevailing belief that joint training for the two tasks leads to performance degradation, we demonstrate that VSSFlow maintains superior performance during end-to-end joint learning process. Furthermore, we use a straightforward feature-level data synthesis method, demonstrating that our framework provides a robust foundation that easily adapts to joint sound and speech generation using synthetic data. Extensive experiments on V2S, VisualTTS and joint generation benchmarks show that VSSFlow effectively unifies these tasks and surpasses state-of-the-art domain-specific baselines, underscoring the critical potential of unified generative models. Project page: https://vasflow1.github.io/vasflow/

音视频生成统一模型扩散模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。