让文字精准控制视频配乐,实现音画同步与创意定制。
CAFA: a Controllable Automatic Foley Artist

- 用文本+视频联合驱动音频生成,通过模态适配器融合多源信息。
- 在语义对齐和音画同步上优于现有方法,文本控制效果显著。
- 适合需要创意配音或个性化音效的视频创作者使用。
Foley 是视频制作中的关键环节,指为无声视频添加音频信号,确保语义与时间上的对齐。近年来,个性化内容创作兴起及视频转音频模型的发展,推动了对用户控制力的需求。现有方法虽支持文本引导音频生成,但在跨模态兼容性方面仍存挑战,尤其当文本引入额外信息或与视觉线索矛盾时。本文提出 CAFA(Controllable Automatic Foley Artist),一个基于文本与视频生成音频的模型,可生成与视频语义和时间严格对齐的音频。该模型以文本到音频模型为基础,通过模态适配器整合视频信息。用户可通过文本细化语义细节或引入创意变化,超越仅依赖视觉线索的生成。实验表明,该方法在语义对齐与音画同步方面表现优异,且在主观与客观评估中均展现出高文本可控性。
原文摘要 · Abstract (English)
Foley is a key element in video production, refers to the process of adding an audio signal to a silent video while ensuring semantic and temporal alignment. In recent years, the rise of personalized content creation and advancements in automatic video-to-audio models have increased the demand for greater user control in the process. One possible approach is to incorporate text to guide audio generation. While supported by existing methods, challenges remain in ensuring compatibility between modalities, particularly when the text introduces additional information or contradicts the sounds naturally inferred from the visuals. In this work, we introduce CAFA (Controllable Automatic Foley Artist) a video-and-text-to-audio model that generates semantically and temporally aligned audio for a given video, guided by text input. CAFA is built upon a text-to-audio model and integrates video information through a modality adapter mechanism. By incorporating text, users can refine semantic details and introduce creative variations, guiding the audio synthesis beyond the expected video contextual cues. Experiments show that besides its superior quality in terms of semantic alignment and audio-visual synchronization the proposed method enable high textual controllability as demonstrated in subjective and objective evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。