arXiv:2503.10719cs.CV2025-03被引 8

用多智能体协作实现长视频配音,提升音画同步与跨场景一致性。

Long-Video Audio Synthesis with Multi-Agent Collaboration

  • 分四步协同处理:分镜、脚本生成、声效设计、音频合成
  • 在207段专业长视频上实现优于基线方法的音画对齐效果
  • 适合影视后期、交互媒体领域研究者与从业者参考

视频到音频合成能为视觉内容生成同步音频,显著提升电影和互动媒体中的沉浸感与叙事连贯性。然而,长视频配音仍面临动态语义变化、时间错位及缺乏专用数据集等挑战。现有方法在短视频上表现良好,但在长视频(如电影)中易出现合成碎片化与跨场景不一致问题。本文提出LVAS-Agent,一种模拟专业配音流程的多智能体框架,将长视频合成分解为四个步骤:场景分割、剧本生成、声音设计与音频合成。核心创新包括用于场景/剧本优化的讨论-修正机制,以及实现时空语义对齐的生成-检索循环。为系统评估,我们构建了首个基准测试集LVAS-Bench,包含207段涵盖多种场景的专业级长视频。实验表明,该方法在音画对齐方面显著优于基线模型。

原文摘要 · Abstract (English)

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an unsolved challenge due to dynamic semantic shifts, temporal misalignment, and the absence of dedicated datasets. While existing methods excel in short videos, they falter in long scenarios (e.g., movies) due to fragmented synthesis and inadequate cross-scene consistency. We propose LVAS-Agent, a novel multi-agent framework that emulates professional dubbing workflows through collaborative role specialization. Our approach decomposes long-video synthesis into four steps including scene segmentation, script generation, sound design and audio synthesis. Central innovations include a discussion-correction mechanism for scene/script refinement and a generation-retrieval loop for temporal-semantic alignment. To enable systematic evaluation, we introduce LVAS-Bench, the first benchmark with 207 professionally curated long videos spanning diverse scenarios. Experiments demonstrate superior audio-visual alignment over baseline methods. Project page: https://lvas-agent.github.io

音频合成多智能体长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。