用户画图就能控制角色骨骼生成与变形,支持从零开始设计和精细修改。
ViP-Rig: Visual-Prompted Controllable Rigging

- 用2D草图和刚度图作为视觉提示,注入预训练模型控制骨骼结构和变形。
- 在Articulation-XL2.0上比几何引导方法更准确恢复目标骨骼与权重。
- 适合需要精准控制角色动画的艺术家或动画师使用。
绑定(Rigging)本质上是任务依赖的,同一网格在不同动画任务中可能需要不同的骨骼结构和变形行为。实践中,艺术家常需反复调整初始绑定的骨骼与变形方式以满足需求。现有自动方法主要基于几何生成合理绑定,但对骨骼和变形行为的显式控制有限。本文提出ViP-Rig,一种视觉提示驱动的绑定框架,支持提示优先生成与结果引导编辑。通过将用户绘制或修改的2D骨骼草图与刚度图特征,注入冻结的预训练骨干网络实现控制。该框架分为两阶段:第一阶段,使用密集到紧凑的视觉提示编码处理骨骼草图,生成固定长度的条件标记,通过门控适配器注入自回归生成器,控制关节位置与分支结构,同时保留生成器的几何先验;第二阶段,采用相同编码设计处理刚度图,冻结的皮肤绑定骨干保持不变,将生成标记对称注入点与关节流,调节点-关节兼容性与皮肤权重。在Articulation-XL2.0上的实验及ModelsResource上的零样本评估表明,ViP-Rig在提示引导下比几何条件基线更准确恢复目标骨骼与皮肤权重。定性结果进一步展示了提示优先生成与结果引导编辑中的显式、局部控制能力。
原文摘要 · Abstract (English)
Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice, artists often inspect an initial rig and repeatedly edit its skeletal structure and deformation behavior to meet specific animation requirements. Existing automatic methods primarily generate a plausible rig from geometry, offering limited explicit control over the resulting skeleton and deformation behavior. In this work, we present ViP-Rig, a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones. Specifically, ViP-Rig consists of two stages, Skeleton Generation and Skinning Prediction. In the first stage, the skeletal sketch is processed by the Dense-to-Compact Visual Prompt Encoding to produce compact, fixed-length conditioning tokens. The resulting tokens are injected into a frozen pretrained autoregressive generator through gated adapters to control joint placement and branching structure while preserving the generator's geometric prior. In the second stage, the rigidity map is processed using the same visual encoding design, while the pretrained skinning backbone remains frozen. The resulting tokens are symmetrically injected into the point and joint streams to modulate point-joint compatibility and the resulting skinning weights. Experiments on Articulation-XL2.0 and zero-shot evaluation on ModelsResource show that ViP-Rig more accurately recovers target skeletons and skinning weights than geometry-conditioned baselines under prompt-guided evaluation. Qualitative results further demonstrate explicit and localized control in both prompt-first rigging and result-guided editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。