用新视角合成图像提升机器人抓取精度,无需多相机移动
Exploiting Radiance Fields for Grasp Generation on Novel Synthetic Views
- 用辐射场生成新视角图像,补充真实拍摄的视觉信息
- 在Graspnet-1billion上实现更多闭合力抓取,覆盖范围更广
- 适合做单图抓取、少样本场景下机器人视觉操控的研究者
基于视觉的机器人操作通常依赖摄像头捕捉包含待操作物体的场景图像。多视角拍摄可缓解遮挡问题,但需移动相机,受限于可达性且耗时。而如高斯点云渲染这类场景表示方法,可从用户指定的新视角生成逼真的虚拟图像。本文首次表明,利用新视角合成能为抓取姿态生成提供额外上下文信息。在Graspnet-1billion数据集上的实验显示,相比稀疏真实视图,新视角合成不仅生成了额外的闭合力抓取(force-closure grasp),还提升了整体抓取覆盖率。未来希望将该方法拓展至仅凭单张输入图像构建辐射场,并结合扩散模型或通用辐射场实现更优抓取提取。
原文摘要 · Abstract (English)
Vision based robot manipulation uses cameras to capture one or more images of a scene containing the objects to be manipulated. Taking multiple images can help if any object is occluded from one viewpoint but more visible from another viewpoint. However, the camera has to be moved to a sequence of suitable positions for capturing multiple images, which requires time and may not always be possible, due to reachability constraints. So while additional images can produce more accurate grasp poses due to the extra information available, the time-cost goes up with the number of additional views sampled. Scene representations like Gaussian Splatting are capable of rendering accurate photorealistic virtual images from user-specified novel viewpoints. In this work, we show initial results which indicate that novel view synthesis can provide additional context in generating grasp poses. Our experiments on the Graspnet-1billion dataset show that novel views contributed force-closure grasps in addition to the force-closure grasps obtained from sparsely sampled real views while also improving grasp coverage. In the future we hope this work can be extended to improve grasp extraction from radiance fields constructed with a single input image, using for example diffusion models or generalizable radiance fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。