音频视频联合去噪能提升视频质量,即使只关注视频本身。
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
- 用预训练文本到音视频模块构建联合去噪架构
- 在复杂运动场景下视频质量显著提升
- 适合想提升视频物理合理性的生成研究者
近期音视频生成系统表明,多模态耦合不仅能增强音视频同步性,还能改善视频自身质量。我们提出核心问题:若仅关注视频质量,音视频联合去噪训练是否仍有益?为此,我们设计参数高效的音视频全迪特(AVFullDiT)架构,利用预训练的文本到视频(T2V)与文本到音频(T2A)模块实现联合去噪。在相同设置下训练了(i)T2AV模型与(ii)T2V-only对照模型。结果首次系统证明,音视频联合去噪不仅提升同步性,还显著改善视频质量,尤其在大尺度及物体接触运动的挑战子集上表现更优。我们推测,预测音频作为特权信号,促使模型内化视觉事件与其声学后果间的因果关系(如碰撞×撞击声),从而正则化视频动态。研究提示跨模态协同训练是构建更强、更符合物理规律世界模型的可行路径。代码与数据集将公开。
原文摘要 · Abstract (English)
Recent audio-video generative systems suggest that coupling modalities benefits not only audio-video synchrony but also the video modality itself. We pose a fundamental question: Does audio-video joint denoising training improve video generation, even when we only care about video quality? To study this, we introduce a parameter-efficient Audio-Video Full DiT (AVFullDiT) architecture that leverages pre-trained text-to-video (T2V) and text-to-audio (T2A) modules for joint denoising. We train (i) a T2AV model with AVFullDiT and (ii) a T2V-only counterpart under identical settings. Our results provide the first systematic evidence that audio-video joint denoising can deliver more than synchrony. We observe consistent improvements on challenging subsets featuring large and object contact motions. We hypothesize that predicting audio acts as a privileged signal, encouraging the model to internalize causal relationships between visual events and their acoustic consequences (e.g., collision $\times$ impact sound), which in turn regularizes video dynamics. Our findings suggest that cross-modal co-training is a promising approach to developing stronger, more physically grounded world models. Code and dataset will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。