arXiv:2605.28394cs.CVcs.GR2026-05

用文本控制手绘2D草图生成3D动画,让动作更自然真实。

Sketch2Motion: Text-driven 2D Sketch to 3D Animation via Diffusion-guided Skeleton Optimization

论文配图:Sketch2Motion: Text-driven 2D Sketch to 3D Animation via Diffusion-guided Skeleton Optimization
图 1 · 摘自论文原文
  • 通过骨骼变换结合扩散模型生成动作,无需成对数据。
  • 引入物理约束和弹簧质量模拟,提升动画稳定性与真实感。
  • 支持人形、四足及非生物类角色,适合创意设计与动画制作。

2D手绘草图的动画化是视觉表达的有效方式,但存在遮挡处理难、动作映射不准等问题。尽管3D动画能更好解决这些问题,但3D运动估计仍十分复杂。现有方法多局限于特定运动类型,如双足行走或面部表情。本文提出Sketch2Motion,一种基于骨骼的扩散引导运动合成框架,将传统角色动画流程与深度生成先验结合。该方法以骨骼变换表示运动,并通过线性混合皮肤技术传播至网格变形。为生成符合语义且逼真的动作,引入文本到视频扩散模型,利用运动感知得分蒸馏采样(MoSDS)实现无配对数据优化。同时施加物理启发的平滑性、拓扑与接触约束,确保运动合理性。进一步集成弹簧-质量模拟器,引入次级运动效果。所提框架具有通用性、可微分性、模块化,兼容双足、四足及非生物类关节角色。实验表明,该方法生成的动画时间连贯、与文本对齐,优于缺乏生成先验或显式物理约束的基线方法。代码与数据集将公开。

原文摘要 · Abstract (English)

Animation of 2D hand-drawn sketches provides an effective medium for visual communication. However, these sketches pose challenges, particularly in handling occlusions and accurately mapping motion. While 3D animation naturally addresses these challenges, estimating 3D motion remains a very complex task. Recent approaches to converting 2D sketches to 3D animations have mainly focused on specific types of motion, such as bipedal movements and facial expressions. We propose Sketch2Motion, a diffusion-guided framework for skeleton-based motion synthesis that combines classical character animation pipelines with deep generative priors. Our method represents motion using skeletal transformations, which are propagated to mesh deformations via linear blend skinning. To guide the resulting animation toward realistic and semantically meaningful motion, we integrate a text-to-video diffusion model via motion-aware score-distillation sampling (MoSDS), enabling optimization without paired motion data. Additionally, we apply physics-inspired smoothness, topological, and contact constraints to stabilize optimization and preserve motion plausibility. Further, we integrate a spring-mass simulator to introduce secondary motion effects. The proposed framework is generalized, fully differentiable, modular, and compatible with biped, quadruped, and non-living articulated characters. Experiments demonstrate that our approach produces temporally coherent, text-aligned animations that outperform baseline motion transfer methods that lack generative priors or explicit physical constraints. We will make our code and dataset publicly available.

2D转3D动画生成扩散模型骨骼驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。