让机器人自动识别物体几何特征并关联使用功能,提升操作智能性。
PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- 通过几何特征聚合自动提取关键点与轴线,跨类别识别物体结构。
- 利用视觉语言模型动态关联几何特征与任务功能描述,实现语义锚定。
- 无需人工标注即可达到近似手动标注效果,适合复杂场景机器人应用。
高阶任务语义与低阶几何特征之间的割裂仍是机器人操作中的核心难题。尽管视觉语言模型(VLMs)在生成具有功能感知的视觉表征方面展现出潜力,但其在标准空间中缺乏语义锚定且依赖人工标注,严重限制了对动态语义-功能关系的捕捉能力。为此,我们提出原语感知语义锚定(PASG)框架,包含:(1) 基于几何特征聚合的自动原语提取,实现跨类别关键点与轴线检测;(2) 由VLM驱动的语义锚定机制,动态关联几何原语与功能属性及任务相关描述;(3) 一个空间-语义推理基准与微调后的VLM(Qwen2.5VL-PA)。我们在多种实际操作场景中验证了PASG的有效性,性能接近人工标注水平。该方法实现了对物体更细粒度的语义-功能理解,建立了一个统一范式,用于在机器人操作中连接几何原语与任务语义。
原文摘要 · Abstract (English)
The fragmentation between high-level task semantics and low-level geometric features remains a persistent challenge in robotic manipulation. While vision-language models (VLMs) have shown promise in generating affordance-aware visual representations, the lack of semantic grounding in canonical spaces and reliance on manual annotations severely limit their ability to capture dynamic semantic-affordance relationships. To address these, we propose Primitive-Aware Semantic Grounding (PASG), a closed-loop framework that introduces: (1) Automatic primitive extraction through geometric feature aggregation, enabling cross-category detection of keypoints and axes; (2) VLM-driven semantic anchoring that dynamically couples geometric primitives with functional affordances and task-relevant description; (3) A spatial-semantic reasoning benchmark and a fine-tuned VLM (Qwen2.5VL-PA). We demonstrate PASG's effectiveness in practical robotic manipulation tasks across diverse scenarios, achieving performance comparable to manual annotations. PASG achieves a finer-grained semantic-affordance understanding of objects, establishing a unified paradigm for bridging geometric primitives with task semantics in robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。