arXiv:2509.11125cs.ROcs.CV2025-09中稿 · RA-L被引 9

让机器人在不同视角下都能稳定抓取,靠的是3D空间解耦表征。

ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations

  • 通过自监督解耦学习,构建视角不变的3D特征表示。
  • 在10个仿真和5个真实任务中,成功率比现有方法高40.6%。
  • 仅用80%参数量,就能实现强鲁棒性和跨域泛化能力。

将视觉强化学习(RL)部署于真实世界操作任务时,常受摄像头视角变化影响。固定视角训练的策略在视角偏移时会失效——这在传感器位置难以精确调控的真实场景中不可避免。现有方法依赖精确标定或难以应对大幅视角变化。为此,我们提出ManiVID-3D,一种面向机器人操作的新型3D RL架构,通过自监督解耦特征学习实现视角不变表征。该框架引入ViewNet模块,无需外参标定即可自动将任意视角下的点云数据对齐至统一空间坐标系。同时开发高效GPU加速批量渲染模块,可实现每秒处理超5000帧,显著提升3D视觉强化学习的大规模训练速度。在10个仿真任务与5个真实任务上的评估表明,该方法在视角变化下成功率比当前最优方法高出40.6%,且参数量减少80%。系统对严重透视变化具有强鲁棒性,并展现出优异的模拟到现实迁移性能,验证了几何一致性表征在非结构化环境中的可扩展性。

原文摘要 · Abstract (English)

Deploying visual reinforcement learning (RL) policies in real-world manipulation is often hindered by camera viewpoint changes. A policy trained from a fixed front-facing camera may fail when the camera is shifted -- an unavoidable situation in real-world settings where sensor placement is hard to manage appropriately. Existing methods often rely on precise camera calibration or struggle with large perspective changes. To address these limitations, we propose ManiVID-3D, a novel 3D RL architecture designed for robotic manipulation, which learns view-invariant representations through self-supervised disentangled feature learning. The framework incorporates ViewNet, a lightweight yet effective module that automatically aligns point cloud observations from arbitrary viewpoints into a unified spatial coordinate system without the need for extrinsic calibration. Additionally, we develop an efficient GPU-accelerated batch rendering module capable of processing over 5000 frames per second, enabling large-scale training for 3D visual RL at unprecedented speeds. Extensive evaluation across 10 simulated and 5 real-world tasks demonstrates that our approach achieves a 40.6% higher success rate than state-of-the-art methods under viewpoint variations while using 80% fewer parameters. The system's robustness to severe perspective changes and strong sim-to-real performance highlight the effectiveness of learning geometrically consistent representations for scalable robotic manipulation in unstructured environments.

机器人操作3D表征强化学习视角不变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。