arXiv:2603.10616cs.RO2026-03

让机器人自适应决定清障或直接抓取,提升复杂环境下的抓握成功率。

AdaClearGrasp: Learning Adaptive Clearing for Zero-Shot Robust Dexterous Grasping in Densely Cluttered Environments

  • 通过视觉语言模型判断是否需清障,统一调用基础操作技能。
  • 零样本泛化抓握成功率在密集杂乱环境中提升显著,最高达87.5%。
  • 适合需要智能决策的机器人抓取场景,尤其适用于复杂环境任务。

在密集杂乱环境中,物理干扰、视觉遮挡和不稳定接触常导致直接灵巧抓取失败,而激进的分离策略可能影响安全。因此,让机器人自适应判断是否清障或直接抓取至关重要。本文提出AdaClearGrasp,一种闭环决策-执行框架,实现密集杂乱环境中的零样本灵巧抓取与自适应清障。该框架将操作视为可控的高层决策过程,由预训练视觉语言模型(VLM)解析视觉观测与语言指令,推理抓取干扰并生成规划骨架,通过统一动作接口调用结构化原子技能。针对灵巧抓取,采用基于相对手物距离表示的强化学习策略,实现跨不同几何与物理属性的零样本泛化。执行过程中,视觉反馈监控结果并在失败时触发重规划,形成闭环修正机制。为评估语言引导的灵巧抓取在杂乱环境中的表现,我们构建了首个具有分级杂乱程度的仿真基准Clutter-Bench,包含3个杂乱等级、7种目标物体,共210个任务场景,并在3种物体、3种杂乱等级下完成18个真实世界实验。结果表明,AdaClearGrasp显著提升了密集杂乱环境下的抓取成功率。

原文摘要 · Abstract (English)

In densely cluttered environments, physical interference, visual occlusions, and unstable contacts often cause direct dexterous grasping to fail, while aggressive singulation strategies may compromise safety. Enabling robots to adaptively decide whether to clear surrounding objects or directly grasp the target is therefore crucial for robust manipulation. We propose AdaClearGrasp, a closed-loop decision-execution framework for adaptive clearing and zero-shot dexterous grasping in densely cluttered environments. The framework formulates manipulation as a controllable high-level decision process that determines whether to directly grasp the target or first clear surrounding objects. A pretrained vision-language model (VLM) interprets visual observations and language task descriptions to reason about grasp interference and generate a high-level planning skeleton, which invokes structured atomic skills through a unified action interface. For dexterous grasping, we train a reinforcement learning policy with a relative hand-object distance representation, enabling zero-shot generalization across diverse object geometries and physical properties. During execution, visual feedback monitors outcomes and triggers replanning upon failures, forming a closed-loop correction mechanism. To evaluate language-conditioned dexterous grasping in clutter, we introduce Clutter-Bench, the first simulation benchmark with graded clutter complexity. It includes seven target objects across three clutter levels, yielding 210 task scenarios. We further perform sim-to-real experiments on three objects under three clutter levels (18 scenarios). Results demonstrate that AdaClearGrasp significantly improves grasp success rates in densely cluttered environments. For more videos and code, please visit our project website: https://chenzixuan99.github.io/adaclear-grasp.github.io/.

灵巧抓取自适应决策零样本仿真基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。