arXiv:2512.24731cs.CV2025-12被引 4

让声音生成更精准可控,支持事件级细粒度编辑。

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

  • 以事件为中心设计分层控制机制,实现声音何时、何物、如何出现的精确指定。
  • 在6000+三元组数据集上,可控性提升40.7%,感知质量提升12.5%。
  • 适合影视音效创作、交互式内容生成等需要精细声效控制的场景。

音效是多模态叙事的关键组成部分,塑造视频的情感氛围与语义结构。尽管视频-文本到音频(VT2A)技术取得进展,当前方法仍存在三大局限:视觉与文本条件不平衡导致视觉主导;缺乏对细粒度可控生成的明确定义;指令理解能力弱,因现有数据集依赖简短类别标签。为此,我们提出新任务EchoFoley,支持事件级局部控制与分层语义控制。通过符号化表示发声事件,明确指定声音在视频或指令中的时间、内容与方式,实现生成、插入与编辑的细粒度控制。为支撑该任务,我们构建了超6000个视频-指令-标注三元组的专家标注基准集EchoFoley-6k。基于此,我们提出基于事件中心的代理生成框架EchoVidia,采用慢-快思维策略。实验表明,EchoVidia在可控性上超越近期模型40.7%,感知质量提升12.5%。

原文摘要 · Abstract (English)

Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-to-audio (VT2A), the current formulation faces three key limitations: First, an imbalance between visual and textual conditioning that leads to visual dominance; Second, the absence of a concrete definition for fine-grained controllable generation; Third, weak instruction understanding and following, as existing datasets rely on brief categorical tags. To address these limitations, we introduce EchoFoley, a new task designed for video-grounded sound generation with both event level local control and hierarchical semantic control. Our symbolic representation for sounding events specifies when, what, and how each sound is produced within a video or instruction, enabling fine-grained controls like sound generation, insertion, and editing. To support this task, we construct EchoFoley-6k, a large-scale, expert-curated benchmark containing over 6,000 video-instruction-annotation triplets. Building upon this foundation, we propose EchoVidia a sounding-event-centric agentic generation framework with slow-fast thinking strategy. Experiments show that EchoVidia surpasses recent VT2A models by 40.7% in controllability and 12.5% in perceptual quality.

声音生成事件控制多模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。