arXiv:2601.16212cs.RO2026-01被引 10

用点云表示打通仿真与现实,仅靠合成数据就能让机器人零样本迁移到真实世界

Point Bridge: 3D Representations for Cross Domain Policy Learning

  • 用视觉语言模型自动提取点云表示,不依赖视觉或物体对齐
  • 零样本迁移最高提升44%,加少量真机演示再增66%性能
  • 适合做仿真到现实迁移的机器人学习研究者

机器人基础模型正逐步实现通用机器人的愿景,但大规模真实世界操作数据的匮乏仍限制进展。仿真和合成数据生成提供了可扩展的替代方案,但其有效性受限于仿真与现实之间的视觉域差距。本文提出 Point Bridge 框架,利用统一、领域无关的点基表示,实现无需显式视觉或物体级对齐的零样本仿真到现实策略迁移。该框架结合基于视觉语言模型(VLMs)的自动化点云表示提取、基于Transformer的策略学习以及高效的推理时管道,仅使用合成数据即可训练出具备真实世界操作能力的智能体。通过在少量真实示范数据上进行联合训练,性能进一步提升,在单任务与多任务设置下,零样本迁移性能最高提升44%,加入少量真实数据后最高提升66%。视频展示请访问:https://pointbridge3d.github.io/

原文摘要 · Abstract (English)

Robot foundation models are beginning to deliver on the promise of generalist robotic agents, yet progress remains constrained by the scarcity of large-scale real-world manipulation datasets. Simulation and synthetic data generation offer a scalable alternative, but their usefulness is limited by the visual domain gap between simulation and reality. In this work, we present Point Bridge, a framework that leverages unified, domain-agnostic point-based representations to unlock synthetic datasets for zero-shot sim-to-real policy transfer, without explicit visual or object-level alignment. Point Bridge combines automated point-based representation extraction via Vision-Language Models (VLMs), transformer-based policy learning, and efficient inference-time pipelines to train capable real-world manipulation agents using only synthetic data. With additional co-training on small sets of real demonstrations, Point Bridge further improves performance, substantially outperforming prior vision-based sim-and-real co-training methods. It achieves up to 44% gains in zero-shot sim-to-real transfer and up to 66% with limited real data across both single-task and multitask settings. Videos of the robot are best viewed at: https://pointbridge3d.github.io/

机器人学习仿真到现实点云表示VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。