用视频生成模型打造更生动的动画分镜,提升角色一致性和故事还原度。
AnimeAgent: Is the Multi-Agent via Image-to-Video models a Good Disney Storytelling Artist?
- 基于图像转视频模型,利用隐式运动先验增强动态表现力
- 在多个数据集上实现最高的一致性、提示遵循度和风格化水平
- 适合动画创作、影视分镜生成等需要高表达力的场景
定制分镜生成(CSG)旨在生成高质量、多角色一致的叙事内容。当前基于静态扩散模型的方法,无论是一次性推理还是多智能体框架,存在三大局限:(1) 静态模型缺乏动态表现力,常出现“复制粘贴”式重复;(2) 一次性推理无法迭代修正缺失属性或提示偏离;(3) 多智能体依赖不鲁棒的评估器,难以评判风格化、非写实动画。为此,我们提出AnimeAgent,首个基于图像转视频(I2V)的多智能体框架用于CSG。受迪士尼“直行动画与定格动画结合”工作流程启发,AnimeAgent利用I2V的隐式运动先验提升一致性与表现力,同时引入混合主观-客观评审机制,实现可靠迭代优化。我们还构建了一个人工标注的CSG基准数据集,包含真实标签。实验表明,AnimeAgent在一致性、提示保真度和风格化方面均达到当前最优性能。
原文摘要 · Abstract (English)
Custom Storyboard Generation (CSG) aims to produce high-quality, multi-character consistent storytelling. Current approaches based on static diffusion models, whether used in a one-shot manner or within multi-agent frameworks, face three key limitations: (1) Static models lack dynamic expressiveness and often resort to "copy-paste" pattern. (2) One-shot inference cannot iteratively correct missing attributes or poor prompt adherence. (3) Multi-agents rely on non-robust evaluators, ill-suited for assessing stylized, non-realistic animation. To address these, we propose AnimeAgent, the first Image-to-Video (I2V)-based multi-agent framework for CSG. Inspired by Disney's "Combination of Straight Ahead and Pose to Pose" workflow, AnimeAgent leverages I2V's implicit motion prior to enhance consistency and expressiveness, while a mixed subjective-objective reviewer enables reliable iterative refinement. We also collect a human-annotated CSG benchmark with ground-truth. Experiments show AnimeAgent achieves SOTA performance in consistency, prompt fidelity, and stylization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。