让视频配音违背视觉内容,实现反事实音效生成
CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

- 分两阶段采样:先构建时间结构并压制视觉音源,再专注塑造目标音色
- 在反事实音效生成上显著优于负提示和现有最优方法
- 适合音效设计、影视后期等需要创意音效的场景
我们研究反事实视频音效生成,即在无声视频基础上,生成与视觉证据矛盾但时间同步的声音。现有视频-文本到音频(VT2A)模型在此任务中常受视觉暗示束缚,难以摆脱视觉所暗示的声音来源。本文提出 ConterFlow,一种针对预训练流匹配型 VT2A 模型的推理时双阶段采样方案。第一阶段基于视频构建时间结构,同时抑制视觉暗示的声音源;第二阶段移除视频条件,专注于将音频音色调整至目标提示。ConterFlow 在反事实音效生成上显著优于朴素负提示及当前最优基线。为评估替代质量,我们提出一种基于文本-音频联合嵌入空间的度量方法,可同时衡量目标提示的支持程度与残留视觉音源泄露。视频演示与代码已公开于 https://gyubin-lee.github.io/counterflow-demo/
原文摘要 · Abstract (English)
We investigate Counterfactual Video Foley Generation, which aims to adopt a sound-source identity that contradicts the visual evidence while remaining temporally synchronized to a silent video. Existing Video&Text-to-Audio (VT2A) models struggle with this, often remaining anchored to the visually implied sound source when video and text contents disagree. We present ConterFlow, an inference-time dual-phase sampling scheme for pretrained flow-matching VT2A models. Phase 1 builds a video-derived temporal structure while suppressing the visually implied source; Phase 2 drops video conditioning to focus entirely on shaping audio timbre toward the target prompt. ConterFlow substantially improves counterfactual Video Foley generation compared to naive negative prompting and state-of-the-art baselines. To evaluate replacement quality, we propose a metric leveraging a text-audio co-embedding space to measure both target-prompt evidence and residual visually implied source leakage. Video demonstrations and code are available at https://gyubin-lee.github.io/counterflow-demo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。