arXiv:2601.08871cs.SDcs.AI2026-01被引 1

用视觉语义线索提升音频混音,实现电影级自动配乐。

Semantic visually-guided acoustic highlighting with large vision-language models

  • 用大视觉语言模型提取六类视觉语义特征
  • 镜头焦点、氛围和场景背景提升效果最显著
  • 为自动化高质量音效设计提供可行路径

平衡对话、音乐与音效并配合视频,对沉浸式叙事至关重要,但当前音频混音流程仍以人工为主且耗时。尽管近期提出了视觉引导的音频强调任务,即利用多模态提示隐式重平衡音频源,但尚不清楚哪些视觉特征最适合作为条件信号。本文通过系统研究深度视频理解是否能改善音频重混,使用文本描述作为视觉分析的代理,提示大视觉语言模型提取六类视觉-语义特征:物体与角色外观、情绪、镜头焦点、色调、场景背景以及推断出的声音相关线索。大量实验表明,镜头焦点、色调和场景背景在感知混音质量上相较现有最优基线表现最佳。研究结果(i)识别出最能支持连贯且视觉对齐音频重混的视觉-语义线索;(ii)提出了一条基于轻量级大视觉语言模型引导的自动化电影级音效设计实用路径。

原文摘要 · Abstract (English)

Balancing dialogue, music, and sound effects with accompanying video is crucial for immersive storytelling, yet current audio mixing workflows remain largely manual and labor-intensive. While recent advancements have introduced the visually guided acoustic highlighting task, which implicitly rebalances audio sources using multimodal guidance, it remains unclear which visual aspects are most effective as conditioning signals.We address this gap through a systematic study of whether deep video understanding improves audio remixing. Using textual descriptions as a proxy for visual analysis, we prompt large vision-language models to extract six types of visual-semantic aspects, including object and character appearance, emotion, camera focus, tone, scene background, and inferred sound-related cues. Through extensive experiments, camera focus, tone, and scene background consistently yield the largest improvements in perceptual mix quality over state-of-the-art baselines. Our findings (i) identify which visual-semantic cues most strongly support coherent and visually aligned audio remixing, and (ii) outline a practical path toward automating cinema-grade sound design using lightweight guidance derived from large vision-language models.

音频生成视觉语言模型音效设计多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。