让声音自动对齐画面动作,还能按意图调节音色。
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
- 分两阶段生成:先提取运动节奏,再按语义生成声音
- 生成的声音与画面动作严格同步,且符合用户指定的材质/动作类型
- 适合影视音效师快速生成专业级拟音,提升创作效率
传统拟音工作依赖手动将音频与画面动作对齐,耗时且难以扩展。尽管视觉到音频生成有进展,但实现时间上连贯、语义可控的声音仍具挑战。为此,我们提出FolAI,一种两阶段生成框架,将声音合成中的‘何时’与‘何物’分离:第一阶段从视频中估计平滑控制信号,捕捉运动强度与节奏结构,作为音频的时间骨架;第二阶段采用基于扩散模型的生成器,根据该时间包络和用户提供的高层语义嵌入(如材质或动作类型)生成声音。这种模块化设计使时间与音色均可精准控制,在保持创意灵活性的同时简化重复性任务。在多种视觉场景(如脚步声生成、动作特异性声效)上的结果表明,模型能可靠生成与画面运动同步、语义一致且听感真实的音频。这验证了FolAI在专业与交互场景下可扩展、高质量拟音合成的潜力。补充材料见 https://ispamm.github.io/FolAI。
原文摘要 · Abstract (English)
Traditional sound design workflows rely on manual alignment of audio events to visual cues, as in Foley sound design, where everyday actions like footsteps or object interactions are recreated to match the on-screen motion. This process is time-consuming, difficult to scale, and lacks automation tools that preserve creative intent. Despite recent advances in vision-to-audio generation, producing temporally coherent and semantically controllable sound effects from video remains a major challenge. To address these limitations, we introduce FolAI, a two-stage generative framework that decouples the when and the what of sound synthesis, i.e., the temporal structure extraction and the semantically guided generation, respectively. In the first stage, we estimate a smooth control signal from the video that captures the motion intensity and rhythmic structure over time, serving as a temporal scaffold for the audio. In the second stage, a diffusion-based generative model produces sound effects conditioned both on this temporal envelope and on high-level semantic embeddings, provided by the user, that define the desired auditory content (e.g., material or action type). This modular design enables precise control over both timing and timbre, streamlining repetitive tasks while preserving creative flexibility in professional Foley workflows. Results on diverse visual contexts, such as footstep generation and action-specific sonorization, demonstrate that our model reliably produces audio that is temporally aligned with visual motion, semantically consistent with user intent, and perceptually realistic. These findings highlight the potential of FolAI as a controllable and modular solution for scalable, high-quality Foley sound synthesis in professional and interactive settings. Supplementary materials are accessible on our dedicated demo page at https://ispamm.github.io/FolAI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。