arXiv:2606.19958cs.CV2026-06

用少量草图和参考图生成结构可控的动画,省时省力。

SketchKeyAnime: Reference-anchored Sparse Key-Sketch Animation Synthesis

论文配图:SketchKeyAnime: Reference-anchored Sparse Key-Sketch Animation Synthesis
图 1 · 摘自论文原文
  • 通过双分支条件机制融合草图与参考图信息
  • 相比最佳基线,草图保真度提升31.9%,时间一致性更好
  • 适合低成本、高控制力的动画创作场景

传统动画制作高度依赖手工绘制和反复修改,尤其在关键帧设计、中间帧补全和角色上色方面。现有动画与视频生成方法通常依赖RGB边界帧、密集帧条件或完整草图序列,难以适应低成本输入。我们提出SketchKeyAnime,一种基于视频扩散框架的方法,可从稀疏关键草图输入生成结构可控、外观一致且时间连贯的动画。给定一张参考RGB图像和若干时序标注的关键草图,SketchKeyAnime引入双分支条件机制,编码局部几何约束与语义-时序上下文。其利用草图交叉注意力融合参考图与草图条件,并引入可学习门控机制;同时采用自适应加权损失强化对关键草图帧和线稿区域的监督。在Sakuga-42M数据集的美学子集上的实验表明,该方法持续优于代表性动画插值与草图引导生成基线。相较于最佳基线,SketchKeyAnime将EDMD降低31.9%,FVD降低9.5%,显著提升草图保真度与时间一致性,多数定量指标表现最优。结果验证了所提框架的有效性,凸显其在低成本、高可控动画创作中的潜力。

原文摘要 · Abstract (English)

Traditional animation production relies heavily on manual drawing and iterative refinement, particularly for key-pose design, in-betweening, and character coloring. While existing animation and video generation methods have made notable progress, they typically depend on RGB boundary frames, dense frame-wise conditions, or complete sketch sequences, limiting their applicability under low-cost input conditions. We present SketchKeyAnime, a video diffusion framework for generating structurally controllable, appearance-consistent, and temporally coherent animations from sparse key-sketch inputs. Given a single reference RGB image and a few temporally indexed key sketches, SketchKeyAnime introduces a dual-branch conditioning mechanism to encode local geometric constraints alongside semantic-temporal context. It leverages Sketch Cross Attention to fuse reference image and sketch conditions with learnable gating, and incorporates an Adaptive Weighted Loss to strengthen supervision on key-sketch frames and line-art regions. Experimental results on the Aesthetic subset of Sakuga-42M show that our approach consistently outperforms representative animation interpolation and sketch-guided generation baselines. Compared to the best-performing baseline, SketchKeyAnime reduces EDMD by 31.9\% and FVD by 9.5\%, demonstrating superior sketch fidelity and temporal coherence, while achieving the best overall performance across most quantitative metrics. These results validate the proposed framework and highlight its potential for low-cost, highly controllable animation creation.

动画生成草图控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。