arXiv:2606.29473cs.CV2026-06中稿 · ECCV

实现多镜头音视频叙事生成,支持自定义剧情控制。

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

论文配图:MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
图 1 · 摘自论文原文
  • 通过边界感知注意力对齐音视频时间边界。
  • 利用身份感知传播保持角色视觉与音色一致性。
  • 构建多智能体脚本系统,将自由输入转为分层描述。

现有生成模型在复杂叙事控制下难以实现连贯的多镜头音视频生成,存在时间错位、可控性差和脚本不完整等问题。本文提出 MAVIN,首个支持自定义叙事控制的多镜头音视频生成框架。为解决时间错位,提出边界感知注意力,结合分层字幕与边界感知标记路由,确保音视频元素在对应时间段内呈现。为提升多主体场景的可控性,提出身份感知传播机制,利用身份嵌入与身份感知掩码绑定特定身份的视觉外观与声线特征。同时构建多智能体脚本管道,将用户自由输入转化为分层字幕。此外,构建了 MAVINSet 数据集,用于模型训练与评估。大量实验表明,MAVIN 达到当前最优性能,为生成模型融入专业影视制作流程开辟新路径。

原文摘要 · Abstract (English)

While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-visual generation. Existing methods suffer from temporal misalignment, limited controllability, and incomplete scripting. In this paper, we propose MAVIN, the first framework for multi-shot audio-visual generation with customized narrative control. To resolve temporal misalignment, we propose boundary-aware attention, which leverages hierarchical captions and boundary-aware token routing to render audio-visual elements within their respective temporal boundaries. To improve the controllability for multi-subject scenarios, we propose ID-aware propagation, utilizing identity embeddings and an identity-aware mask to bind specific identities to consistent visual appearances and vocal timbres. To provide comprehensive audio-visual narratives, we present a multi-agent scripting pipeline to transform free-form user inputs into hierarchical captions. Furthermore, we construct MAVINSet, a multi-shot audio-visual dataset for robust training and evaluation. Extensive experiments demonstrate that MAVIN achieves state-of-the-art performance, opening up a new avenue for integrating generative models into professional filmmaking workflows.

音视频生成叙事控制多镜头

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。