arXiv:2505.12154cs.CVcs.SD2025-05CVPR被引 5

让音频随视频画面自动突出重点,提升视听一致性。

Learning to Highlight Audio by Watching Movies

  • 用视觉信息指导音频强化,通过多模态Transformer实现
  • 在新数据集上显著优于基线,主观评价更自然流畅
  • 适合影视音效制作、跨模态生成研究者参考

近年来视频内容创作与消费激增,但音频与视觉的焦点匹配仍显不足。本文提出视觉引导的音频突出任务,旨在根据视频内容动态调整音频重点,增强视听协调性。为此设计了一种基于Transformer的多模态框架,并构建了名为Muddy Mix的新数据集,利用电影中精细的音视频制作提供自由监督信号。通过分离、调节、重混三步流程模拟真实场景下的劣质混音,训练模型在复杂条件下仍能有效提取听觉焦点。实验表明,该方法在定量与主观评估中均优于多个基线模型,且系统分析了不同上下文引导方式及数据难度的影响。

原文摘要 · Abstract (English)

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal viewpoint selection or post-editing, has been central to media production, its natural counterpart, audio, has not undergone equivalent advancements. This often results in a disconnect between visual and acoustic saliency. To bridge this gap, we introduce a novel task: visually-guided acoustic highlighting, which aims to transform audio to deliver appropriate highlighting effects guided by the accompanying video, ultimately creating a more harmonious audio-visual experience. We propose a flexible, transformer-based multimodal framework to solve this task. To train our model, we also introduce a new dataset -- the muddy mix dataset, leveraging the meticulous audio and video crafting found in movies, which provides a form of free supervision. We develop a pseudo-data generation process to simulate poorly mixed audio, mimicking real-world scenarios through a three-step process -- separation, adjustment, and remixing. Our approach consistently outperforms several baselines in both quantitative and subjective evaluation. We also systematically study the impact of different types of contextual guidance and difficulty levels of the dataset. Our project page is here: https://wikichao.github.io/VisAH/.

音频生成多模态视听一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。