用机器人动作对齐训练稀疏3D表示,实现跨任务通用感知
Sparse2Act: Learning Action-Aligned Sparse 3D Representations for Cross-Domain Robot Manipulation

- 以末端执行器动作作为几何监督,让点云编码器对齐操作空间
- 在LIBERO-10上微调500步后成功率86.9%,跨域迁移达73.4%
- 支持真实世界部署,仿真预训练+少量实数据微调即达72.5%成功率
显式3D表示因能暴露物体形状、工作空间几何及机器人-物体关系,在操作任务中具有吸引力。然而,稀疏3D编码器通常通过下游任务目标学习,导致表示依赖特定数据分布、策略架构和动作参数化。本文提出Sparse2Act,一种用于预训练稀疏点云编码器的观察-动作对齐框架。核心思想是使用任务空间末端执行器动作作为几何监督:对掩码稀疏3D标记进行训练,使其围绕与观测配对的工作空间运动组织场景特征。预训练后,仅复用编码器初始化,下游策略可保留原有架构和动作空间(包括关节空间命令)。在LIBERO-10基准上,微调500步后平均成功率达86.9%。同一预训练编码器支持LIBERO到Meta-World跨域迁移,在Meta-World-5上平均成功率达73.4%。消融实验表明,性能提升来自掩码动作对齐信号,且对下游动作解码器容量不敏感。真实世界实验显示,仿真预训练后结合有限真实数据微调,四类任务平均成功率为72.5%,证明了有效的模拟到现实迁移能力。结果表明,机器人动作可为可复用的稀疏3D表示提供紧凑的几何监督。
原文摘要 · Abstract (English)
Explicit 3D representations are attractive for manipulation because they expose object shape, workspace geometry, and robot-object relations in metric coordinates. However, sparse 3D encoders are often learned through downstream task objectives, tying the representation to a particular data distribution, policy architecture, and action parameterization. We introduce Sparse2Act, an observation-action alignment framework for pretraining sparse point-cloud encoders. The key idea is to use task-space end-effector actions as geometric supervision: masked sparse 3D tokens are trained to organize scene features around the workspace motion paired with the observation. After pretraining, only the encoder initialization is reused by downstream policies, allowing them to retain their own architectures and action spaces, including joint-space commands. On the LIBERO-10 benchmark, our method achieves 86.9% average success after 500 fine-tuning steps. The same pretrained encoder supports LIBERO-to-Meta-World cross-domain transfer, achieving 73.4% average success on the Meta-World-5 benchmark. Ablations on the objective and decoder capacity show that the gains come from the masked action-alignment signal and remain useful across downstream action decoders. In real-world experiments, simulation pretraining followed by limited real-data fine-tuning achieves an average success rate of 72.5% across four tasks, demonstrating effective sim-to-real transfer. These results suggest that robot actions can provide compact geometric supervision for reusable sparse 3D representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。