arXiv:2510.03117cs.CVcs.SD2025-10被引 12

通过分离图文描述与双向交互机制,实现文本生成音视频的精准同步。

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

  • 采用分层视觉引导的双描述框架,分离视频与音频的文本条件。
  • 提出双塔扩散变压器,实现语义与时间上的音视频同步,性能领先。
  • 适合关注跨模态生成、音视频对齐的研究者与开发者。

本研究聚焦于文本到音视频生成(T2SV)这一挑战性任务,旨在从文本条件生成同步的音视频内容,并确保两个模态均与文本一致。尽管联合音视频训练已取得进展,但仍有两大难题未解决:(1) 使用单一共享文本描述时,视频与音频的文本相同会引发模态干扰,混淆预训练模型;(2) 跨模态特征交互的最优机制仍不明确。为此,我们提出分层视觉引导的双描述框架(HVGC),生成解耦的视频与音频描述,消除条件阶段的干扰。在此基础上,引入BridgeDiT——一种新型双塔扩散变换器,采用双交叉注意力(DCA)机制,作为稳健的“桥梁”,实现对称且双向的信息交换,达成语义与时间上的同步。在三个基准数据集上的大量实验,结合人工评估,表明该方法在多数指标上达到当前最优。全面消融实验进一步验证了各模块的有效性,为未来T2SV任务提供了关键洞见。所有代码与模型权重将公开发布。

原文摘要 · Abstract (English)

This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanwhile ensuring both modalities are aligned with text. Despite progress in joint audio-video training, two critical challenges still remain unaddressed: (1) a single, shared text caption where the text for video is equal to the text for audio often creates modal interference, confusing the pretrained backbones, and (2) the optimal mechanism for cross-modal feature interaction remains unclear. To address these challenges, we first propose the Hierarchical Visual-Grounded Captioning (HVGC) framework that generates pairs of disentangled captions, a video caption, and an audio caption, eliminating interference at the conditioning stage. Based on HVGC, we further introduce BridgeDiT, a novel dual-tower diffusion transformer, which employs a Dual CrossAttention (DCA) mechanism that acts as a robust ``bridge" to enable a symmetric, bidirectional exchange of information, achieving both semantic and temporal synchronization. Extensive experiments on three benchmark datasets, supported by human evaluations, demonstrate that our method achieves state-of-the-art results on most metrics. Comprehensive ablation studies further validate the effectiveness of our contributions, offering key insights for the future T2SV task. All the codes and checkpoints will be publicly released.

音视频生成跨模态扩散模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。