arXiv:2602.03028cs.CV2026-02被引 6

用多智能体闭环机制让长视频故事更连贯一致。

MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration

  • 通过计划-执行-验证-修订循环,动态控制角色、场景和时间线。
  • 在长序列生成中显著降低语义漂移,保持角色一致性。
  • 适合需要高连贯性的影视生成、互动叙事等应用。

从简短用户提示生成长篇音视频故事仍具挑战,源于意图与执行间的鸿沟:需在长时间跨度内保持高层叙事意图的一致性。现有方法多依赖前馈流程或仅靠提示优化,导致语义漂移和身份不一致。本文提出MUSE,一种基于闭环约束强化的多智能体框架,通过迭代的计划-执行-验证-修订循环协调生成过程。MUSE将叙事意图转化为可机器执行的显式控制,涵盖角色身份、空间构图与时间连续性,并在生成中施加针对性多模态反馈以纠正偏差。为评估无真实参考的开放式叙事,我们引入MUSEBench,一种经人工判断验证的无参考评价协议。实验表明,MUSE相比代表性基线显著提升长时序叙事连贯性、跨模态身份一致性及电影级质量。

原文摘要 · Abstract (English)

Generating long-form audio-visual stories from a short user prompt remains challenging due to an intent-execution gap, where high-level narrative intent must be preserved across coherent, shot-level multimodal generation over long horizons. Existing approaches typically rely on feed-forward pipelines or prompt-only refinement, which often leads to semantic drift and identity inconsistency as sequences grow longer. We address this challenge by formulating storytelling as a closed-loop constraint enforcement problem and propose MUSE, a multi-agent framework that coordinates generation through an iterative plan-execute-verify-revise loop. MUSE translates narrative intent into explicit, machine-executable controls over identity, spatial composition, and temporal continuity, and applies targeted multimodal feedback to correct violations during generation. To evaluate open-ended storytelling without ground-truth references, we introduce MUSEBench, a reference-free evaluation protocol validated by human judgments. Experiments demonstrate that MUSE substantially improves long-horizon narrative coherence, cross-modal identity consistency, and cinematic quality compared with representative baselines.

多智能体故事生成长视频闭环控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。