无需标注数据,用因果表示学习解耦机器人姿态的生成因素。
ROPES: Robotic Pose Estimation via Score-Based Causal Representation Learning
- 基于得分的因果表示学习,自动识别可操控的生成因子。
- 在半合成机械臂实验中实现高保真度因子解耦,无需标签数据。
- 适合对无监督学习、机器人感知感兴趣的学者与工程师。
因果表示学习(CRL)作为一种强大的无监督框架,能够解耦高维数据背后的生成因素,并学习这些解耦变量间的因果关系。尽管近年来在可辨识性和实际应用方面取得进展,理论与现实实践之间仍存在显著差距。本文通过将CRL引入机器人领域,迈出缩小这一差距的关键一步。具体而言,针对明确的机器人位姿估计任务——从原始图像中恢复位置与姿态——提出了基于得分的因果表示学习的机器人位姿估计方法(ROPES)。作为无监督框架,ROPES体现干预式因果学习的核心思想:识别被动作控制的生成因素。图像由内在和外在潜在因素(如关节角度、手臂/肢体几何、光照、背景及相机配置)生成,目标是解耦并恢复可直接操控(干预)的潜在变量。根据干预式因果学习理论,通过干预产生变化的变量可被识别。在机器人中,这可通过命令各关节执行器并记录不同控制下的图像自然实现。半合成机械臂实验表明,ROPES能以高保真度解耦潜在生成因素,且仅依赖分布变化,无需任何标注数据。论文还与一种近期提出的半监督基线进行了对比。最后,本文将机器人位姿估计定位为因果表示学习的近实用测试平台。
原文摘要 · Abstract (English)
Causal representation learning (CRL) has emerged as a powerful unsupervised framework that (i) disentangles the latent generative factors underlying high-dimensional data, and (ii) learns the cause-and-effect interactions among the disentangled variables. Despite extensive recent advances in identifiability and some practical progress, a substantial gap remains between theory and real-world practice. This paper takes a step toward closing that gap by bringing CRL to robotics, a domain that has motivated CRL. Specifically, this paper addresses the well-defined robot pose estimation -- the recovery of position and orientation from raw images -- by introducing Robotic Pose Estimation via Score-Based CRL (ROPES). Being an unsupervised framework, ROPES embodies the essence of interventional CRL by identifying those generative factors that are actuated: images are generated by intrinsic and extrinsic latent factors (e.g., joint angles, arm/limb geometry, lighting, background, and camera configuration) and the objective is to disentangle and recover the controllable latent variables, i.e., those that can be directly manipulated (intervened upon) through actuation. Interventional CRL theory shows that variables that undergo variations via interventions can be identified. In robotics, such interventions arise naturally by commanding actuators of various joints and recording images under varied controls. Empirical evaluations in semi-synthetic manipulator experiments demonstrate that ROPES successfully disentangles latent generative factors with high fidelity with respect to the ground truth. Crucially, this is achieved by leveraging only distributional changes, without using any labeled data. The paper also includes a comparison with a baseline based on a recently proposed semi-supervised framework. This paper concludes by positioning robot pose estimation as a near-practical testbed for CRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。