arXiv:2411.04925cs.CVcs.AI2024-11被引 29

多智能体协作生成定制化叙事视频,角色一致性显著提升。

StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration

  • 分角色智能体分工协作,模拟专业影视制作流程
  • 自研LoRA-BE方法提升单镜头内时序一致性,跨镜头保持主角一致
  • 适合需要高角色连贯性的个性化视频生成场景

人工智能生成内容(AIGC)推动了自动化视频生成研究,但定制化叙事视频的生成仍面临角色一致性难维持的挑战。现有方法如Mora和AesopAgent虽采用多智能体架构进行故事到视频(S2V)生成,但在主角一致性与定制化支持方面表现不足。为此,我们提出StoryAgent,一种面向定制化叙事视频生成(CSVG)的多智能体框架。该框架将任务分解为故事设计、分镜生成、视频创作、智能体协调与结果评估等环节,由专用智能体协同完成。通过融合不同模型优势,显著增强生成过程控制力。我们引入定制化图像到视频(I2V)方法LoRA-BE,提升单镜头内时序一致性;并设计新型分镜生成管道,保障跨镜头主体一致性。大量实验表明,本方法在生成高度一致的叙事视频方面优于当前最优方法。贡献包括StoryAgent框架及提升主角一致性的创新技术。

原文摘要 · Abstract (English)

The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video production, particularly for customized narratives, remains challenging due to the complexity of maintaining subject consistency across shots. While existing approaches like Mora and AesopAgent integrate multiple agents for Story-to-Video (S2V) generation, they fall short in preserving protagonist consistency and supporting Customized Storytelling Video Generation (CSVG). To address these limitations, we propose StoryAgent, a multi-agent framework designed for CSVG. StoryAgent decomposes CSVG into distinct subtasks assigned to specialized agents, mirroring the professional production process. Notably, our framework includes agents for story design, storyboard generation, video creation, agent coordination, and result evaluation. Leveraging the strengths of different models, StoryAgent enhances control over the generation process, significantly improving character consistency. Specifically, we introduce a customized Image-to-Video (I2V) method, LoRA-BE, to enhance intra-shot temporal consistency, while a novel storyboard generation pipeline is proposed to maintain subject consistency across shots. Extensive experiments demonstrate the effectiveness of our approach in synthesizing highly consistent storytelling videos, outperforming state-of-the-art methods. Our contributions include the introduction of StoryAgent, a versatile framework for video generation tasks, and novel techniques for preserving protagonist consistency.

视频生成多智能体角色一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。