arXiv:2604.24842cs.AIcs.MA2026-04被引 1

用多智能体框架让视频生成更连贯,避免故事跑偏。

Co-Director: Agentic Generative Video Storytelling

论文配图:Co-Director: Agentic Generative Video Storytelling
图 1 · 摘自论文原文
  • 分层智能体设计:全局选创意方向,局部优化细节一致性。
  • 在400个虚构产品广告场景中,显著超越现有方法。
  • 适合想构建连贯影视叙事的开发者与研究者。

尽管扩散模型能生成高保真视频片段,但将其转化为连贯的故事创作引擎仍具挑战。现有代理流水线通过串联模块自动化实现,却因独立且手工设计的提示词导致语义漂移和级联失败。本文提出Co-Director,一种分层多智能体框架,将视频叙事建模为全局优化问题。为确保语义连贯,引入分层参数化:全局多臂赌博机识别有前景的创意方向,局部多模态自精炼循环缓解身份漂移,保障序列一致性。该设计平衡了新颖叙事策略的探索与有效创作配置的利用。为评估,构建GenAD-Bench数据集,包含400个虚构产品个性化广告场景。实验表明,Co-Director显著优于当前最先进基线,提供可推广至更广泛电影叙事的系统性方案。

原文摘要 · Abstract (English)

While diffusion models generate high-fidelity video clips, transforming them into coherent storytelling engines remains challenging. Current agentic pipelines automate this via chained modules but suffer from semantic drift and cascading failures due to independent, handcrafted prompting. We present Co-Director, a hierarchical multi-agent framework formalizing video storytelling as a global optimization problem. To ensure semantic coherence, we introduce hierarchical parameterization: a multi-armed bandit globally identifies promising creative directions, while a local multimodal self-refinement loop mitigates identity drift and ensures sequence-level consistency. This balances the exploration of novel narrative strategies with the exploitation of effective creative configurations. For evaluation, we introduce GenAD-Bench, a 400-scenario dataset of fictional products for personalized advertising. Experiments demonstrate that Co-Director significantly outperforms state-of-the-art baselines, offering a principled approach that seamlessly generalizes to broader cinematic narratives. Project Page: https://co-director-agent.github.io/

视频生成多智能体叙事一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。