arXiv:2510.24261cs.ROcs.AI2025-10NeurIPS被引 2

通过未来渲染学习3D动态,提升机器人操作的泛化能力

DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation

  • 用可微体素渲染实现掩码重建与未来预测,联合学习3D几何、语义和动态
  • 在RLBench和Colosseum上成功率显著提升,对环境扰动有强鲁棒性
  • 适合需要3D理解与动态预测的机器人抓取、放置等任务

由于真实世界训练数据多样性的匮乏,学习通用的机器人操作策略仍是关键挑战。现有方法多依赖2D视觉自监督预训练(如掩码图像建模),主要关注静态语义或场景几何;或使用大规模视频预测模型,强调2D动态,难以同时捕捉操作所需的几何、语义与动态信息。本文提出DynaRend,一种基于可微体素渲染的表征学习框架,通过掩码重建与未来预测,学习3D感知且富含动态信息的三平面特征。在多视角RGB-D视频数据上预训练后,该框架统一捕获空间几何、未来动态与任务语义。所学表征可通过动作价值图预测有效迁移至下游机器人操作任务。在RLBench和Colosseum两个基准上以及真实机器人实验中均验证了其优势,显著提升了策略成功率、对环境扰动的泛化能力及实际应用效果。

原文摘要 · Abstract (English)

Learning generalizable robotic manipulation policies remains a key challenge due to the scarcity of diverse real-world training data. While recent approaches have attempted to mitigate this through self-supervised representation learning, most either rely on 2D vision pretraining paradigms such as masked image modeling, which primarily focus on static semantics or scene geometry, or utilize large-scale video prediction models that emphasize 2D dynamics, thus failing to jointly learn the geometry, semantics, and dynamics required for effective manipulation. In this paper, we present DynaRend, a representation learning framework that learns 3D-aware and dynamics-informed triplane features via masked reconstruction and future prediction using differentiable volumetric rendering. By pretraining on multi-view RGB-D video data, DynaRend jointly captures spatial geometry, future dynamics, and task semantics in a unified triplane representation. The learned representations can be effectively transferred to downstream robotic manipulation tasks via action value map prediction. We evaluate DynaRend on two challenging benchmarks, RLBench and Colosseum, as well as in real-world robotic experiments, demonstrating substantial improvements in policy success rate, generalization to environmental perturbations, and real-world applicability across diverse manipulation tasks.

机器人操作3D表征自监督学习动态预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。