arXiv:2507.02271cs.CVcs.AI2025-07IJCAI被引 4

让视频生成音频模型学会识别部分可见画面中的电影语言。

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation

  • 用自蒸馏模拟电影语言变化,让模型理解局部画面与声音的关联。
  • 在部分可见场景下,所有评估指标均显著提升。
  • 适合需要精准音效匹配的影视后期制作人员。

视频到音频(V2A)生成已取得显著进展,在影视后期制作中扮演关键角色。然而,现有方法忽视了电影语言这一重要的艺术表达元素,导致在福莱目标仅部分可见时性能下降。为此,我们提出一种简单的自蒸馏方法,将V2A模型扩展至电影语言场景。通过模拟电影语言的变化,学生模型学习对齐训练样本中视频特征与相同音视频对应关系,从而有效捕捉声音与部分视觉信息之间的关联。该方法不仅在部分可见场景下各项评估指标均有显著提升,还在大规模V2A数据集VGGSound上增强了性能。

原文摘要 · Abstract (English)

Video-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language, a critical component of artistic expression in filmmaking. As a result, their performance deteriorates in scenarios where Foley targets are only partially visible. To address this challenge, we propose a simple self-distillation approach to extend V2A models to cinematic language scenarios. By simulating the cinematic language variations, the student model learns to align the video features of training pairs with the same audio-visual correspondences, enabling it to effectively capture the associations between sounds and partial visual information. Our method not only achieves impressive improvements under partial visibility across all evaluation metrics, but also enhances performance on the large-scale V2A dataset, VGGSound.

视频生成音频生成电影语言自蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。