用生成模型解决音视频焦点不一致问题,让声音更贴合画面。
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
- 将音频重平衡建模为生成任务,用流匹配方法处理模糊映射。
- 引入滚动损失,减少早期错误累积,提升长程生成稳定性。
- 融合音视频特征进行源选择,适合跨模态内容创作场景。
视觉引导的音频强调旨在使音频与伴随视频协调一致,创造连贯的音视频体验。尽管视觉显著性和增强已广泛研究,音频强调仍处于探索阶段,常导致视听焦点错位。现有方法采用判别式模型,难以应对音频重混中的固有歧义——不良平衡与良好平衡音频之间不存在自然的一一对应关系。为此,我们将该任务重新定义为生成问题,提出条件流匹配(Conditional Flow Matching, CFM)框架。迭代流生成的一个关键挑战是:早期预测错误(如选错需增强的声源)会随步骤累积,导致轨迹偏离流形。为此,我们设计了一种滚动损失,在最终步骤惩罚偏移,鼓励自纠正轨迹并稳定长程流集成。此外,我们提出了一个条件模块,在向量场回归前融合音频与视觉线索,实现显式的跨模态声源选择。大量定量和定性评估表明,我们的方法持续优于先前的最优判别式方法,证实视觉引导的音频重混应通过生成建模解决。
原文摘要 · Abstract (English)
Visually-guided acoustic highlighting seeks to rebalance audio in alignment with the accompanying video, creating a coherent audio-visual experience. While visual saliency and enhancement have been widely studied, acoustic highlighting remains underexplored, often leading to misalignment between visual and auditory focus. Existing approaches use discriminative models, which struggle with the inherent ambiguity in audio remixing, where no natural one-to-one mapping exists between poorly-balanced and well-balanced audio mixes. To address this limitation, we reframe this task as a generative problem and introduce a Conditional Flow Matching (CFM) framework. A key challenge in iterative flow-based generation is that early prediction errors -- in selecting the correct source to enhance -- compound over steps and push trajectories off-manifold. To address this, we introduce a rollout loss that penalizes drift at the final step, encouraging self-correcting trajectories and stabilizing long-range flow integration. We further propose a conditioning module that fuses audio and visual cues before vector field regression, enabling explicit cross-modal source selection. Extensive quantitative and qualitative evaluations show that our method consistently surpasses the previous state-of-the-art discriminative approach, establishing that visually-guided audio remixing is best addressed through generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。