arXiv:2507.04959cs.CVcs.AI2025-07被引 1

用户点击视频对象即可生成对应声音,实现精准可控的音视频合成。

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation

  • 通过点击指定物体,结合掩码引导视觉编码器提取对象级特征。
  • 在多个数据集上音视频对应性提升,新指标CAV分数显著优于基线。
  • 适合影视制作、交互式内容创作等需要细粒度音效控制的场景。

视频到音频(V2A)生成在电影制作等领域具有巨大潜力。尽管取得进展,现有方法依赖全局视频信息,在复杂场景中难以生成特定对象的音效。为此,我们提出Hear-Your-Click,一种通过点击视频帧来生成特定对象声音的交互式V2A框架。我们设计了基于掩码引导的视觉编码器(MVE)与面向对象的对比音频-视觉微调(OCAV),以获得与音频对齐的对象级视觉特征。同时,引入两种数据增强策略:随机视频拼接(RVS)和掩码引导音量调节(MLM),提升模型对分割对象的敏感度。为评估音视频对应性,我们设计了新的评价指标CAV分数。大量实验表明,该框架实现了更精确的控制,并在多项指标上提升生成性能。

原文摘要 · Abstract (English)

Video-to-audio (V2A) generation shows great potential in fields such as film production. Despite significant advances, current V2A methods relying on global video information struggle with complex scenes and generating audio tailored to specific objects. To address these limitations, we introduce Hear-Your-Click, an interactive V2A framework enabling users to generate sounds for specific objects by clicking on the frame. To achieve this, we propose Object-aware Contrastive Audio-Visual Fine-tuning (OCAV) with a Mask-guided Visual Encoder (MVE) to obtain object-level visual features aligned with audio. Furthermore, we tailor two data augmentation strategies, Random Video Stitching (RVS) and Mask-guided Loudness Modulation (MLM), to enhance the model's sensitivity to segmented objects. To measure audio-visual correspondence, we designed a new evaluation metric, the CAV score. Extensive experiments demonstrate that our framework offers more precise control and improves generation performance across various metrics. Project Page: https://github.com/SynapGrid/Hear-Your-Click

音视频生成交互式生成对象级控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。