arXiv:2410.13851cs.ROcs.CV2024-10CoRL被引 28

让机器人视觉外观可微,直接从图像像素优化控制参数。

Differentiable Robot Rendering

  • 用运动学感知的可变形模型结合高斯溅射实现可微渲染。
  • 能从图像重建机器人姿态,并通过视觉语言模型控制机器人。
  • 支持任意机器人结构和自由度,为视觉大模型用在机器人打基础。

视觉基础模型在海量视觉数据上训练后,展现出前所未有的开放世界推理与规划能力。将它们应用于机器人任务的关键挑战在于视觉数据与动作数据之间的模态差距。我们提出可微机器人渲染方法,使机器人本体的视觉外观可直接对控制参数求导。该模型融合了运动学感知的可变形模型与高斯溅射,兼容任意机器人形态和自由度。我们在多个应用中验证其能力,包括从图像重建机器人姿态,以及通过视觉语言模型控制机器人。定量与定性结果表明,该可微渲染模型能从像素直接提供有效的梯度用于机器人控制,为视觉基础模型在机器人领域的未来应用奠定基础。

原文摘要 · Abstract (English)

Vision foundation models trained on massive amounts of visual data have shown unprecedented reasoning and planning skills in open-world settings. A key challenge in applying them to robotic tasks is the modality gap between visual data and action data. We introduce differentiable robot rendering, a method allowing the visual appearance of a robot body to be directly differentiable with respect to its control parameters. Our model integrates a kinematics-aware deformable model and Gaussians Splatting and is compatible with any robot form factors and degrees of freedom. We demonstrate its capability and usage in applications including reconstruction of robot poses from images and controlling robots through vision language models. Quantitative and qualitative results show that our differentiable rendering model provides effective gradients for robotic control directly from pixels, setting the foundation for the future applications of vision foundation models in robotics.

可微渲染机器人控制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。