arXiv:2412.10255cs.GRcs.AI2024-12被引 9

专为动画视频生成设计的高效模型与评测体系

AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era

  • 构建1000万+高质量动画数据处理流水线,支持多种生成任务
  • 引入时空掩码模块,实现帧间插补与局部图像引导动画生成
  • 发布948段动画视频评测集,适配动画风格与夸张运动特性

动画在影视产业中日益受到关注。尽管Sora、Kling和CogVideoX等先进视频生成模型在自然视频生成上表现优异,但在处理动画视频时仍存在不足。动画视频生成评估困难,因其具有独特艺术风格、违反物理规律及夸张动作特征。本文提出AniSora系统,涵盖数据处理流水线、可控生成模型和评估基准。基于超过1000万条高质量数据,生成模型集成时空掩码模块,支持图像到视频生成、帧插值和局部图像引导动画等核心功能。同时收集了948个不同类型的动画视频评测集,并设计专门针对动画生成的评估指标。整个项目已公开于https://github.com/bilibili/Index-anisora/tree/main。

原文摘要 · Abstract (English)

Animation has gained significant interest in the recent film and TV industry. Despite the success of advanced video generation models like Sora, Kling, and CogVideoX in generating natural videos, they lack the same effectiveness in handling animation videos. Evaluating animation video generation is also a great challenge due to its unique artist styles, violating the laws of physics and exaggerated motions. In this paper, we present a comprehensive system, AniSora, designed for animation video generation, which includes a data processing pipeline, a controllable generation model, and an evaluation benchmark. Supported by the data processing pipeline with over 10M high-quality data, the generation model incorporates a spatiotemporal mask module to facilitate key animation production functions such as image-to-video generation, frame interpolation, and localized image-guided animation. We also collect an evaluation benchmark of 948 various animation videos, with specifically developed metrics for animation video generation. Our entire project is publicly available on https://github.com/bilibili/Index-anisora/tree/main.

动画生成视频生成模型评测可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。