提出可抗噪的3D点云机器人操作方法,提升真实场景泛化能力。
EquiForm: Noise-Robust SE(3)-Equivariant Policy Learning from 3D Point Clouds
- 通过几何去噪与对比对齐,修复噪声导致的结构偏差
- 模拟任务平均提升17.2%,真实任务提升28.1%
- 适合高噪声、遮挡严重的工业机器人操作场景
基于3D点云的视觉模仿学习已推动机器人操作发展,提供几何感知且外观不变的观测。然而,点云策略仍对传感器噪声、位姿扰动和遮挡引起的伪影高度敏感,破坏几何结构并违背等变性假设,影响泛化性能。现有等变方法主要在神经网络中编码对称性约束,但未显式修正噪声引起的几何偏差或强制学习表征的等变一致性。本文提出EquiForm,一种面向点云操作的抗噪SE(3)等变策略学习框架。EquiForm阐明了噪声引起的几何畸变如何导致观测-动作映射中的等变性偏差,并引入几何去噪模块,在噪声或不完整观测下恢复一致的3D结构。同时,提出对比等变对齐目标,强制表征在刚体变换和噪声扰动下保持一致性。基于这些组件,EquiForm构建灵活的策略学习流程,融合抗噪几何推理与现代生成模型。在16个模拟任务和4个真实世界操作任务上评估,相比最先进方法,模拟任务平均提升17.2%,真实任务提升28.1%,验证了其强抗噪性和空间泛化能力。
原文摘要 · Abstract (English)
Visual imitation learning with 3D point clouds has advanced robotic manipulation by providing geometry-aware, appearance-invariant observations. However, point cloud-based policies remain highly sensitive to sensor noise, pose perturbations, and occlusion-induced artifacts, which distort geometric structure and break the equivariance assumptions required for robust generalization. Existing equivariant approaches primarily encode symmetry constraints into neural architectures, but do not explicitly correct noise-induced geometric deviations or enforce equivariant consistency in learned representations. We introduce EquiForm, a noise-robust SE(3)-equivariant policy learning framework for point cloud-based manipulation. EquiForm formalizes how noise-induced geometric distortions lead to equivariance deviations in observation-to-action mappings, and introduces a geometric denoising module to restore consistent 3D structure under noisy or incomplete observations. In addition, we propose a contrastive equivariant alignment objective that enforces representation consistency under both rigid transformations and noise perturbations. Built upon these components, EquiForm forms a flexible policy learning pipeline that integrates noise-robust geometric reasoning with modern generative models. We evaluate EquiForm on 16 simulated tasks and 4 real-world manipulation tasks across diverse objects and scene layouts. Compared to state-of-the-art point cloud imitation learning methods, EquiForm achieves an average improvement of 17.2% in simulation and 28.1% in real-world experiments, demonstrating strong noise robustness and spatial generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。