让点云目标直接指导抓取旋转,无需额外姿态信息。
Rotation-Aware Point-Cloud Embeddings for Vision-Based In-Hand Reorientation

- 设计旋转感知的点云嵌入,使距离反映真实旋转误差。
- 模型仅用点云和本体感觉,就能完成精细物体翻转。
- 适合无姿态估计、无教师监督的强化学习场景。
点云目标可直接定义灵巧操作中的物体姿态:无需设定特定于物体的坐标系或在测试时估计6D姿态,策略直接接收物体期望的3D几何形状。然而原始点云目标条件对策略学习而言条件不足。当前与目标点云均无序、独立采样且常受可见性影响,其差异混淆了物体旋转与排列、重采样及不稳定的对应结构。因此,现有方法通常在表示外添加结构,如显式姿态输入、相对姿态特征、密集流特征或来自特权教师的蒸馏。本文通过学习一种旋转感知的点云嵌入,其欧氏距离与SO(3)测地误差对齐,使当前-目标对比变为平滑控制信号。模型仅需当前与目标点云嵌入、本体感觉及质心元数据,无需物体姿态、相对姿态、密集流或教师动作监督,即可执行模型无关的强化学习策略。在手内重定向实验中,该接口达到特权状态和蒸馏基线性能,同时避免了测试时计算结构化姿态或流输入的脆弱性。结果表明,当表征本身编码旋转相关几何时,点云目标才真正适用于此任务。此外,我们还发现通用视觉点云预训练不足以支持此类当前-目标比较,因其丢弃任务相关状态,仅保留形状特征。
原文摘要 · Abstract (English)
Point-cloud goals provide a direct way to specify dexterous in-hand reorientation: instead of defining an object-specific pose frame or estimating 6D pose at test time, the policy is given the desired 3D geometry of the object. Yet raw point-cloud goal conditioning is poorly conditioned for policy learning. Current and goal clouds are unordered, independently sampled, and often visibility-dependent, so their discrepancy entangles object rotation with permutation, resampling, and unstable correspondence structure. For this reason, prior point-cloud manipulation methods typically add structure outside the representation itself, such as explicit pose or relative-pose inputs, dense flow features, or distillation from privileged teachers. We close this gap by learning a rotation-aware point-cloud embedding whose Euclidean latent distance is calibrated to the SO(3) geodesic error between object orientations. The resulting representation turns current-goal comparison into a smooth control signal, allowing a model-free RL policy to act from current and goal point-cloud embeddings, proprioception, and centroid metadata, without object pose, relative pose, dense flow, or teacher-action supervision. In in-hand reorientation experiments, this interface matches privileged-state and distillation-based baselines while avoiding brittle test-time computation of structured pose or flow inputs. These results suggest that point-cloud goals become practical for this task when the representation, rather than an external module, encodes the task-relevant geometry of rotation. We also show evidence that generic visual point-cloud pretraining is insufficient for such a current-goal comparison because it discards the task-relevant state and preserves only shape features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。