arXiv:2606.01940cs.CV2026-06

无需标注或模板,单张3D图像就能解耦物体结构与运动

SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation

论文配图:SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
图 1 · 摘自论文原文
  • 用等变自编码器对齐不同实例到统一坐标系
  • 通过循环重建和可学习模板,准确恢复关节参数
  • 适合无标注数据下的机械臂、家具等可动物体分析

现有类别级可动物体姿态估计方法通常依赖密集监督、多帧输入或CAD模板,难以分离几何与运动,且无法显式恢复关节参数。本文提出SCAPO,一种自监督框架,仅需单张RGB-D图像即可估计规范几何、刚性部件分割及关节轴、枢轴与运动状态,无需真值标签或类别特定模型。SCAPO首先利用SE(3)等变向量神经自编码器去除全局姿态,将多样本对齐至共享规范空间;在此空间中,设计关节感知的混合皮肤模块建模部件运动。通过观测形状与规范形状间的循环重建,以及可学习规范模板在跨空间对齐,实现共性类别几何与个体残差形状的解耦。在合成与真实可动物体数据集上的实验表明,SCAPO能恢复一致的部件结构和精确的关节参数,优于所有自监督基线。

原文摘要 · Abstract (English)

Existing methods for category-level object articulation from a single 3D observation often rely on dense supervision, multi-frame inputs, or CAD templates, and still struggle to disentangle geometry from articulation or to recover explicit joint parameters. We propose SCAPO, a self-supervised framework that estimates canonical geometry, rigid part segmentation, and joint pivots, axes, and articulation states from a single RGB-D observation without ground-truth labels or category-specific models. Our SCAPO first uses an SE(3)-equivariant vector-neuron autoencoder to factor out global pose and align diverse instances into a shared canonical space. On this aligned shape, a joint-aware blend-skinning module is then designed to model part motion. We learn this representation through cycle reconstruction between observed and canonical shapes and cross-space alignment with a learnable canonical template that decouples shared category geometry from instance-specific residual shape. Experiments on synthetic and real articulated-object datasets show that our SCAPO recovers consistent part structure and accurate articulation parameters and outperforms all self-supervised baselines.

姿态估计自监督可动物体3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。