首个多层次标注的动漫生成数据集,支持多模态可控动画生成。
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- 构建分层标注的多模态动漫数据集,覆盖多种生成任务
- 包含40万图像转视频、5万全身动作标注等四类数据,总计超47000条样本
- 适合研究动漫生成、跨模态控制与角色动画的开发者使用
由于非人类角色复杂性、风格化动作和精细情绪表达,实现高质量的多模态卡通动画生成极具挑战。真实视频与卡通动画之间存在巨大领域差异,后者通常具有抽象性与夸张动作。同时,由于大规模自动标注困难,公开的多模态卡通数据极为稀缺。为此,我们提出MagicAnime数据集,一个大规模、分层标注、多模态的数据集,支持多种视频生成任务,并附带基准测试。包含40万图像到视频生成样本、5万视频-关键点对用于全身动作标注、1.2万视频到视频面部动画样本,以及2900对视频-音频对用于音频驱动面部动画。此外,我们构建了MagicAnime-Bench基准,涵盖视频驱动面部动画、音频驱动面部动画、图像到视频动画和姿态驱动角色动画四项任务。在四个任务上的综合实验验证了其在高保真度、细粒度和可控性生成方面的有效性。
原文摘要 · Abstract (English)
Generating high-quality cartoon animations multimodal control is challenging due to the complexity of non-human characters, stylistically diverse motions and fine-grained emotions. There is a huge domain gap between real-world videos and cartoon animation, as cartoon animation is usually abstract and has exaggerated motion. Meanwhile, public multimodal cartoon data are extremely scarce due to the difficulty of large-scale automatic annotation processes compared with real-life scenarios. To bridge this gap, We propose the MagicAnime dataset, a large-scale, hierarchically annotated, and multimodal dataset designed to support multiple video generation tasks, along with the benchmarks it includes. Containing 400k video clips for image-to-video generation, 50k pairs of video clips and keypoints for whole-body annotation, 12k pairs of video clips for video-to-video face animation, and 2.9k pairs of video and audio clips for audio-driven face animation. Meanwhile, we also build a set of multi-modal cartoon animation benchmarks, called MagicAnime-Bench, to support the comparisons of different methods in the tasks above. Comprehensive experiments on four tasks, including video-driven face animation, audio-driven face animation, image-to-video animation, and pose-driven character animation, validate its effectiveness in supporting high-fidelity, fine-grained, and controllable generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。