用离散扩散模型提升可动物体姿态估计的精度与鲁棒性
DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
- 将姿态估计转为离散扩散过程,逐步去噪恢复真实姿态
- 引入动态决策机制,平衡真实与噪声分布,提升建模精度
- 基于层级运动结构先验,适合复杂可动物体的6D姿态估计
可动物体姿态估计是具身智能的核心任务。现有方法通常在连续空间中回归姿态,但面临两大挑战:1)搜索空间庞大复杂;2)难以融入内在运动学约束。本文提出DICArt(DIsCrete Diffusion for Articulation Pose Estimation),将姿态估计建模为条件离散扩散过程。不依赖连续域,而是通过学习的逆扩散过程逐步去噪,恢复真实姿态。为提升建模精度,设计灵活的流动决策器,动态决定每个令牌是否去噪或重置,有效平衡真实与噪声分布。同时引入层级运动耦合策略,分层估计各刚性部分姿态,尊重物体运动结构。在合成与真实世界数据集上验证,结果表明该方法性能更优、鲁棒性更强。通过结合离散生成建模与结构先验,为复杂环境下的类别级6D姿态估计提供了新范式。
原文摘要 · Abstract (English)
Articulated object pose estimation is a core task in embodied AI. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this work, we introduce DICArt (DIsCrete Diffusion for Articulation Pose Estimation), a novel framework that formulates pose estimation as a conditional discrete diffusion process. Instead of operating in a continuous domain, DICArt progressively denoises a noisy pose representation through a learned reverse diffusion procedure to recover the GT pose. To improve modeling fidelity, we propose a flexible flow decider that dynamically determines whether each token should be denoised or reset, effectively balancing the real and noise distributions during diffusion. Additionally, we incorporate a hierarchical kinematic coupling strategy, estimating the pose of each rigid part hierarchically to respect the object's kinematic structure. We validate DICArt on both synthetic and real-world datasets. Experimental results demonstrate its superior performance and robustness. By integrating discrete generative modeling with structural priors, DICArt offers a new paradigm for reliable category-level 6D pose estimation in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。