用孩子的一幅画生成连贯有故事感的动画,保持原画风格。
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
- 分步构建:先生成分镜,再用风格传播保持画风一致。
- 通过3D重建与两阶段运动适配,实现自然动作和身份保留。
- 适合儿童创意表达、个性化动画创作,易用性强。
我们提出FairyGen,一个从单幅儿童绘画自动生成叙事性卡通视频的系统,忠实保留其独特艺术风格。不同于以往侧重角色一致性与基础运动的叙事方法,FairyGen显式解耦角色建模与风格化背景生成,并引入电影镜头设计以支持富有表现力且连贯的叙事。给定单个角色草图,首先利用多模态大模型(MLLM)生成包含场景设定、角色动作与摄像机视角的结构化分镜。为保证视觉一致性,引入风格传播适配器,捕捉角色视觉风格并应用于背景,完整保留角色视觉特征的同时合成风格一致的场景。镜头设计模块基于分镜进行帧裁剪与多视角合成,提升画面多样性和电影质感。为实现动画生成,重构角色3D代理以推导物理合理的运动序列,并用于微调基于MMDiT的图像到视频扩散模型。进一步提出两阶段运动定制适配器:第一阶段从时序无序帧中学习外观特征,解耦身份与动作;第二阶段使用时间步偏移策略建模时序动态,冻结身份权重。训练完成后,FairyGen可直接生成与分镜对齐的多样化、连贯视频场景。大量实验表明,该系统生成的动画在风格忠实度、叙事结构与自然运动方面均表现优异,展现了个性化互动叙事动画的巨大潜力。代码将开源于 https://github.com/GVCLab/FairyGen。
原文摘要 · Abstract (English)
We propose FairyGen, an automatic system for generating story-driven cartoon videos from a single child's drawing, while faithfully preserving its unique artistic style. Unlike previous storytelling methods that primarily focus on character consistency and basic motion, FairyGen explicitly disentangles character modeling from stylized background generation and incorporates cinematic shot design to support expressive and coherent storytelling. Given a single character sketch, we first employ an MLLM to generate a structured storyboard with shot-level descriptions that specify environment settings, character actions, and camera perspectives. To ensure visual consistency, we introduce a style propagation adapter that captures the character's visual style and applies it to the background, faithfully retaining the character's full visual identity while synthesizing style-consistent scenes. A shot design module further enhances visual diversity and cinematic quality through frame cropping and multi-view synthesis based on the storyboard. To animate the story, we reconstruct a 3D proxy of the character to derive physically plausible motion sequences, which are then used to fine-tune an MMDiT-based image-to-video diffusion model. We further propose a two-stage motion customization adapter: the first stage learns appearance features from temporally unordered frames, disentangling identity from motion; the second stage models temporal dynamics using a timestep-shift strategy with frozen identity weights. Once trained, FairyGen directly renders diverse and coherent video scenes aligned with the storyboard. Extensive experiments demonstrate that our system produces animations that are stylistically faithful, narratively structured natural motion, highlighting its potential for personalized and engaging story animation. The code will be available at https://github.com/GVCLab/FairyGen
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。