用可微渲染实现零样本场景重建与抓取
Differentiable Inverse Graphics for Zero-shot Scene Reconstruction and Robot Grasping
- 结合神经基础模型与物理可微渲染,从单张RGBD图重建场景
- 无需额外3D数据或测试时采样,实现零样本抓取
- 适合需要少数据、高可解释性的机器人应用
在未知真实环境中有效运作需机器人能估计并交互未见物体。现有最先进模型依赖大量训练数据和测试时样本构建黑箱场景表示。本文提出一种可微神经图形模型,结合神经基础模型与基于物理的可微渲染,实现无需任何额外3D数据或测试时样本的零样本场景重建与机器人抓取。模型通过求解一系列约束优化问题,从单张RGBD图像和边界框中估计物理一致的场景参数,如网格、光照、材质属性及未见物体的6D姿态。我们在标准无模型少样本基准上评估,结果表明其在无模型少样本姿态估计上优于现有算法。此外,通过零样本抓取任务验证了场景重建的准确性。该方法无需依赖大规模数据集或测试时采样,为新环境中的数据高效、可解释、泛化性强的机器人自主提供了路径。
原文摘要 · Abstract (English)
Operating effectively in novel real-world environments requires robotic systems to estimate and interact with previously unseen objects. Current state-of-the-art models address this challenge by using large amounts of training data and test-time samples to build black-box scene representations. In this work, we introduce a differentiable neuro-graphics model that combines neural foundation models with physics-based differentiable rendering to perform zero-shot scene reconstruction and robot grasping without relying on any additional 3D data or test-time samples. Our model solves a series of constrained optimization problems to estimate physically consistent scene parameters, such as meshes, lighting conditions, material properties, and 6D poses of previously unseen objects from a single RGBD image and bounding boxes. We evaluated our approach on standard model-free few-shot benchmarks and demonstrated that it outperforms existing algorithms for model-free few-shot pose estimation. Furthermore, we validated the accuracy of our scene reconstructions by applying our algorithm to a zero-shot grasping task. By enabling zero-shot, physically-consistent scene reconstruction and grasping without reliance on extensive datasets or test-time sampling, our approach offers a pathway towards more data efficient, interpretable and generalizable robot autonomy in novel environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。