arXiv:2510.21581cs.CVcs.SD2025-10被引 4

用轻量桥接让冻结模型对齐视频与音效,实现精准同步。

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

  • 插入紧凑交叉注意力桥,仅训练少量参数
  • 视频引导音效时间与细节,保持提示可控性
  • 适合需要快速迭代的音效生成场景

Foley Control是一种轻量级视频引导音效生成方法,保持预训练单模态模型冻结,仅学习二者间的少量跨注意力桥梁。通过在冻结的Stable Audio Open DiT文本到音频(T2A)模型中插入紧凑的视频交叉注意力模块,将V-JEPA2视频嵌入与文本提示连接,使提示定义全局语义,视频细化时间节奏和局部动态。冻结主干保留强边缘分布(视频;给定文本的音频),而桥梁学习音视频依赖关系以实现同步,无需重训练音频先验。为减少内存并稳定训练,我们在条件化前对视频标记进行池化。在精心构建的视频-音效基准上,Foley Control在时间和语义对齐上表现优异,可训练参数远少于近期多模态系统,同时保持提示可控性和生产友好型模块化(可无端到端重训更换编码器或T2A主干)。尽管聚焦于视频到音效,该桥接设计也可扩展至其他音频模态(如语音)。

原文摘要 · Abstract (English)

Foley Control is a lightweight approach to video-guided Foley that keeps pretrained single-modality models frozen and learns only a small cross-attention bridge between them. We connect V-JEPA2 video embeddings to a frozen Stable Audio Open DiT text-to-audio (T2A) model by inserting compact video cross-attention after the model's existing text cross-attention, so prompts set global semantics while video refines timing and local dynamics. The frozen backbones retain strong marginals (video; audio given text) and the bridge learns the audio-video dependency needed for synchronization -- without retraining the audio prior. To cut memory and stabilize training, we pool video tokens before conditioning. On curated video-audio benchmarks, Foley Control delivers competitive temporal and semantic alignment with far fewer trainable parameters than recent multi-modal systems, while preserving prompt-driven controllability and production-friendly modularity (swap/upgrade encoders or the T2A backbone without end-to-end retraining). Although we focus on Video-to-Foley, the same bridge design can potentially extend to other audio modalities (e.g., speech).

音效生成跨模态对齐轻量模型冻结主干

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。