解决机器人抓取中语义误判与环境动态性难题,实现高鲁棒性开集桌面物体抓取。
CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping
- 异步闭环感知框架分离语义与几何,减少空间幻觉。
- 执行前后状态比对生成文本反馈,成功率达87.0%。
- 自动生成多模态数据,适合复杂场景与真实部署。
桌面物体抓取广泛应用于智能制造、物流和农业。尽管视觉语言模型(VLMs)在机器人操作中展现出巨大潜力,但其在低层抓取任务中的应用仍面临三大挑战:高质量多模态示范数据稀缺、几何定位薄弱导致的空间幻觉,以及动态环境中开放环执行的脆弱性。为此,我们提出闭合回路异步空间感知(CLASP)框架,融合多模态感知、逻辑推理与状态反馈。首先,设计双路径分层感知模块,解耦高层语义意图与几何定位,引导推理输出并生成确定性动作元组,降低空间误判。其次,实现异步闭合环评估器,对比执行前后的状态,提供文本诊断反馈,建立稳健的纠错机制,提升动态环境下的鲁棒性。最后,构建可扩展的多模态数据引擎,无需人工遥控即可从真实与合成场景自动生成高质量空间标注与推理模板。大量实验表明,该方法显著优于现有基线,在整体任务中达到87.0%的成功率。尤其在多样化物体、几何复杂类别及杂乱场景中表现卓越,有效弥合仿真到现实的差距。
原文摘要 · Abstract (English)
Robot grasping of desktop object is widely used in intelligent manufacturing, logistics, and agriculture.Although vision-language models (VLMs) show strong potential for robotic manipulation, their deployment in low-level grasping faces key challenges: scarce high-quality multimodal demonstrations, spatial hallucination caused by weak geometric grounding, and the fragility of open-loop execution in dynamic environments. To address these challenges, we propose Closed-Loop Asynchronous Spatial Perception(CLASP), a novel asynchronous closed-loop framework that integrates multimodal perception, logical reasoning, and state-reflective feedback. First, we design a Dual-Pathway Hierarchical Perception module that decouples high-level semantic intent from geometric grounding. The design guides the output of the inference model and the definite action tuples, reducing spatial illusions. Second, an Asynchronous Closed-Loop Evaluator is implemented to compare pre- and post-execution states, providing text-based diagnostic feedback to establish a robust error-correction loop and improving the vulnerability of traditional open-loop execution in dynamic environments. Finally, we design a scalable multi-modal data engine that automatically synthesizes high-quality spatial annotations and reasoning templates from real and synthetic scenes without human teleoperation. Extensive experiments demonstrate that our approach significantly outperforms existing baselines, achieving an 87.0% overall success rate. Notably, the proposed framework exhibits remarkable generalization across diverse objects, bridging the sim-to-real gap and providing exceptional robustness in geometrically challenging categories and cluttered scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。