让视频生成音频更可控,能精准应对视觉与文本冲突。
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- 用视觉-音频联合编码提升图文对齐与控制能力。
- 分离时间与音色信息,实现风格精确调控。
- 适合需要高精度音频生成的影视、游戏开发场景。
视频到音频(V2A)生成近年取得进展,但实现稳健且细粒度的可控性仍具挑战。现有方法在视觉与文本冲突下文本控制能力弱,且因参考音频中时间与音色信息纠缠,导致风格控制不精准。此外,缺乏标准化评估基准限制了系统性评测。本文提出ControlFoley,一种统一的多模态V2A框架,支持对视频、文本和参考音频的精细控制。引入融合CLIP与时空音视频编码器的联合视觉编码机制,增强对齐与文本可控性;提出时序-音色解耦策略,在抑制冗余时间线索的同时保留关键音色特征;设计模态鲁棒训练方案,包含统一多模态表示对齐(REPA)与随机模态丢弃。同时构建VGGSound-TVC基准,用于评估不同视觉-文本冲突程度下的文本可控性。大量实验表明,ControlFoley在多类V2A任务中表现领先,包括文本引导、文本控制与音频控制生成。其在跨模态冲突下仍保持优异可控性,同步性与音质表现突出,性能优于或媲美工业级V2A系统。代码、模型、数据集及演示已开源:https://github.com/xiaomi-research/controlfoley。
原文摘要 · Abstract (English)
Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual-text conflict and imprecise stylistic control due to entangled temporal and timbre information in reference audio. Moreover, the lack of standardized benchmarks limits systematic evaluation. We propose ControlFoley, a unified multimodal V2A framework that enables precise control over video, text, and reference audio. We introduce a joint visual encoding paradigm that integrates CLIP with a spatio-temporal audio-visual encoder to improve alignment and textual controllability. We further propose temporal-timbre decoupling to suppress redundant temporal cues while preserving discriminative timbre features. In addition, we design a modality-robust training scheme with unified multimodal representation alignment (REPA) and random modality dropout. We also present VGGSound-TVC, a benchmark for evaluating textual controllability under varying degrees of visual-text conflict. Extensive experiments demonstrate state-of-the-art performance across multiple V2A tasks, including text-guided, text-controlled, and audio-controlled generation. ControlFoley achieves superior controllability under cross-modal conflict while maintaining strong synchronization and audio quality, and shows competitive or better performance compared to an industrial V2A system. Code, models, datasets, and demos are available at: https://github.com/xiaomi-research/controlfoley.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。