构建首个参考引导的多镜头动画数据集,助力连贯动画生成。
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
- 通过自动化流程实现跨镜头视觉一致性与分层标注
- 基于参考图和上下文信息生成视频,显著提升一致性
- 适合动画生成、多模态模型研究者使用
近年来,AI生成内容(AIGC)显著加速了动画制作。为生成具有叙事性的动画,需生成连贯的多镜头视频片段,并结合脚本与角色参考图。然而现有公开数据集主要聚焦真实场景,缺乏角色参考图像以保证视觉一致性。为此,我们提出 AnimeShooter——一个参考引导的多镜头动画数据集。该数据集包含完整分层标注,通过自动化流程保障镜头间强视觉一致性。故事级标注涵盖剧情概要、关键场景及含参考图的角色档案;镜头级标注将故事分解为连续镜头,每镜头附带场景、角色及叙事性与描述性双重视觉标题。此外,子集 AnimeShooter-audio 提供每镜头同步音频、音频描述与声源信息。为验证数据集有效性并建立基线,我们提出 AnimeShooterGen,融合多模态大语言模型(MLLM)与视频扩散模型。参考图与前序生成镜头由 MLLM 处理,生成兼顾参考与上下文的表征,再作为条件输入扩散模型生成下一镜头。实验表明,基于 AnimeShooter 训练的模型在跨镜头视觉一致性和参考图像遵循度上均表现更优,凸显该数据集在连贯动画生成中的价值。
原文摘要 · Abstract (English)
Recent advances in AI-generated content (AIGC) have significantly accelerated animation production. To produce engaging animations, it is essential to generate coherent multi-shot video clips with narrative scripts and character references. However, existing public datasets primarily focus on real-world scenarios with global descriptions, and lack reference images for consistent character guidance. To bridge this gap, we present AnimeShooter, a reference-guided multi-shot animation dataset. AnimeShooter features comprehensive hierarchical annotations and strong visual consistency across shots through an automated pipeline. Story-level annotations provide an overview of the narrative, including the storyline, key scenes, and main character profiles with reference images, while shot-level annotations decompose the story into consecutive shots, each annotated with scene, characters, and both narrative and descriptive visual captions. Additionally, a dedicated subset, AnimeShooter-audio, offers synchronized audio tracks for each shot, along with audio descriptions and sound sources. To demonstrate the effectiveness of AnimeShooter and establish a baseline for the reference-guided multi-shot video generation task, we introduce AnimeShooterGen, which leverages Multimodal Large Language Models (MLLMs) and video diffusion models. The reference image and previously generated shots are first processed by MLLM to produce representations aware of both reference and context, which are then used as the condition for the diffusion model to decode the subsequent shot. Experimental results show that the model trained on AnimeShooter achieves superior cross-shot visual consistency and adherence to reference visual guidance, which highlight the value of our dataset for coherent animated video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。