无需深度相机,仅用摄像头实现复杂环境下的稳定抓取。
GraspView: Active Perception Scoring and Best-View Optimization for Robotic Grasping in Cluttered Environments
- 单张RGB图重建全局3D场景,融合多视角信息保持几何一致。
- 动态选择最佳观测视角,有效解决遮挡问题,提升感知覆盖。
- 适配透明/反光物体,适合真实杂乱场景中的机器人操作。
机器人抓取是自主操作的基础能力,但在存在遮挡、感知质量差和3D重建不一致的杂乱环境中仍具挑战性。传统方法依赖RGB-D相机获取几何信息,对透明或高光物体失效,近距离时性能下降。本文提出GraspView,一种纯RGB的机器人抓取方案,在无深度传感器条件下实现复杂环境中的精准操作。其框架包含三个核心模块:(i) 全局感知场景重建,从单张RGB图像生成局部一致、可度量的3D结构,并融合多视角投影构建连贯全局场景;(ii) 渲染-评分主动感知策略,动态选择下一最佳视角以揭示被遮挡区域;(iii) 在线度量对齐模块,将VGGT预测结果与机器人运动学校准,确保物理尺度一致性。基于这些定制模块,GraspView实现多视角融合的最优视图抓取,结合GraspNet完成鲁棒执行。在多种桌面物体上的实验表明,该方法显著优于基于RGB-D和单视角RGB的基线模型,尤其在严重遮挡、近距传感及透明物体场景下表现突出。结果表明,GraspView为非结构化现实环境中的可靠抓取提供了实用且通用的替代方案。
原文摘要 · Abstract (English)
Robotic grasping is a fundamental capability for autonomous manipulation, yet remains highly challenging in cluttered environments where occlusion, poor perception quality, and inconsistent 3D reconstructions often lead to unstable or failed grasps. Conventional pipelines have widely relied on RGB-D cameras to provide geometric information, which fail on transparent or glossy objects and degrade at close range. We present GraspView, an RGB-only robotic grasping pipeline that achieves accurate manipulation in cluttered environments without depth sensors. Our framework integrates three key components: (i) global perception scene reconstruction, which provides locally consistent, up-to-scale geometry from a single RGB view and fuses multi-view projections into a coherent global 3D scene; (ii) a render-and-score active perception strategy, which dynamically selects next-best-views to reveal occluded regions; and (iii) an online metric alignment module that calibrates VGGT predictions against robot kinematics to ensure physical scale consistency. Building on these tailor-designed modules, GraspView performs best-view global grasping, fusing multi-view reconstructions and leveraging GraspNet for robust execution. Experiments on diverse tabletop objects demonstrate that GraspView significantly outperforms both RGB-D and single-view RGB baselines, especially under heavy occlusion, near-field sensing, and with transparent objects. These results highlight GraspView as a practical and versatile alternative to RGB-D pipelines, enabling reliable grasping in unstructured real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。