用多智能体协作生成长视频,模拟电影制作流程提升创意产出质量。
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
- 构建分层图结构框架OmniAgent,模仿电影制作实现模块化分工。
- 引入超图节点支持临时群组讨论,降低单个智能体记忆负担。
- 采用有限重试的有向环图,支持反馈迭代优化前期生成结果。
近期多智能体系统在提升创意任务表现方面展现出巨大潜力,如长视频生成。本研究提出三项创新:首先,提出OmniAgent——一种基于图结构的分层多智能体框架,借鉴电影制作流程,实现模块化专业化与可扩展的跨智能体协作;其次,受上下文工程启发,引入超图节点,使缺乏充分上下文的智能体能进行临时群体讨论,降低个体记忆需求同时保障上下文完整性;第三,将有向无环图(DAG)替换为带有限重试机制的有向环图,使智能体可通过后续节点反馈对前期输出进行反思与优化,从而提升整体生成质量。这些贡献为创意任务中更稳健的多智能体系统发展奠定基础。
原文摘要 · Abstract (English)
Recent advancements in multi-agent systems have demonstrated significant potential for enhancing creative task performance, such as long video generation. This study introduces three innovations to improve multi-agent collaboration. First, we propose OmniAgent, a hierarchical, graph-based multi-agent framework for long video generation that leverages a film-production-inspired architecture to enable modular specialization and scalable inter-agent collaboration. Second, inspired by context engineering, we propose hypergraph nodes that enable temporary group discussions among agents lacking sufficient context, reducing individual memory requirements while ensuring adequate contextual information. Third, we transition from directed acyclic graphs (DAGs) to directed cyclic graphs with limited retries, allowing agents to reflect and refine outputs iteratively, thereby improving earlier stages through feedback from subsequent nodes. These contributions lay the groundwork for developing more robust multi-agent systems in creative tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。