arXiv:2603.15597cs.SDcs.CV2026-03中稿 · ICLR被引 3

用参考音频精准控制视频生成声音,解决文本描述模糊问题。

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

  • 直接用参考音频条件生成,避免文本描述歧义
  • 实现音色迁移、零样本生成,音频质量优于现有方法
  • 适合需要精细声音控制的影视音效创作场景

现有视频到音频(V2A)生成方法主要依赖文本提示和视觉信息合成音频。但存在两大瓶颈:训练数据中语义粒度不足,如将声学上不同的声音归为粗略标签;以及文本描述难以准确表达微小声学特征。这导致基于文本控制的细粒度声音合成困难。为此,我们提出AC-Foley,一种以参考音频为条件的V2A模型,直接利用参考音频实现对生成声音的精确与细粒度控制。该方法支持细粒度声音合成、音色迁移、零样本声音生成及提升音频质量。通过直接以音频信号为条件,规避了文本描述的语义模糊性,同时实现对声学属性的精确操控。实验证明,当以参考音频为条件时,AC-Foley在福莱音效生成任务上达到当前最优性能;即使无音频条件,仍保持与先进V2A方法相当的竞争力。代码与演示见:https://ff2416.github.io/AC-Foley-Page。

原文摘要 · Abstract (English)

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating acoustically distinct sounds under coarse labels, and textual ambiguity in describing micro-acoustic features. These bottlenecks make it difficult to perform fine-grained sound synthesis using text-controlled modes. To address these limitations, we propose AC-Foley, an audio-conditioned V2A model that directly leverages reference audio to achieve precise and fine-grained control over generated sounds. This approach enables fine-grained sound synthesis, timbre transfer, zero-shot sound generation, and improved audio quality. By directly conditioning on audio signals, our approach bypasses the semantic ambiguities of text descriptions while enabling precise manipulation of acoustic attributes. Empirically, AC-Foley achieves state-of-the-art performance for Foley generation when conditioned on reference audio, while remaining competitive with state-of-the-art video-to-audio methods even without audio conditioning. Code and demo are available at: https://ff2416.github.io/AC-Foley-Page

音视频生成参考音频细粒度控制音色迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。