让视频生成音频时能精准控制声音出现的时间点,像导演一样安排音效。
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
- 用结构化脚本提供分段时间指令,实现细粒度时间控制。
- 在复杂多事件场景下仍保持高保真音频质量,音质与原模型一致。
- 适合需要精确音效调度的影视制作、游戏开发等场景。
近期视频转音频(V2A)方法已取得显著进展,可合成高质量真实音频。然而,在多事件场景或视觉线索不足(如小区域、屏幕外声音、遮挡或部分可见物体)时,难以实现精细的时间控制。本文提出FoleyDirector框架,首次在基于DiT的V2A生成中实现精确的时间引导,同时保持基础模型的音频质量,并支持V2A生成与时间可控合成间的无缝切换。FoleyDirector引入结构化时间脚本(STS),通过短时间片段对应的描述提供更丰富的时序信息。这些特征通过脚本引导的时间融合模块(Script-Guided Temporal Fusion Module)结合,利用时间脚本注意力机制实现协同融合。为应对复杂多事件场景,进一步提出双帧音频合成(Bi-Frame Sound Synthesis),支持帧内与帧外音频并行生成,提升可控性。为支持训练与评估,构建了DirectorSound数据集,并引入VGGSoundDirector和DirectorBench。实验表明,FoleyDirector显著增强时间可控性,同时保持高音频保真度,使用户可如同音效导演般操控生成过程,推动V2A向更富表现力和可控的方向发展。
原文摘要 · Abstract (English)
Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they struggle with fine-grained temporal control in multi-event scenarios or when visual cues are insufficient, such as small regions, off-screen sounds, or occluded or partially visible objects. In this paper, we propose FoleyDirector, a framework that, for the first time, enables precise temporal guidance in DiT-based V2A generation while preserving the base model's audio quality and allowing seamless switching between V2A generation and temporally controlled synthesis. FoleyDirector introduces Structured Temporal Scripts (STS), a set of captions corresponding to short temporal segments, to provide richer temporal information. These features are integrated via the Script-Guided Temporal Fusion Module, which employs Temporal Script Attention to fuse STS features coherently. To handle complex multi-event scenarios, we further propose Bi-Frame Sound Synthesis, enabling parallel in-frame and out-of-frame audio generation and improving controllability. To support training and evaluation, we construct the DirectorSound dataset and introduce VGGSoundDirector and DirectorBench. Experiments demonstrate that FoleyDirector substantially enhances temporal controllability while maintaining high audio fidelity, empowering users to act as Foley directors and advancing V2A toward more expressive and controllable generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。