让机器人理解物体用途并选对可拿的物品,结果透明可解释。
CRAFT-E: A Neuro-Symbolic Framework for Embodied Affordance Grounding
- 用知识图谱+视觉语言对齐,构建可解释的物体功能推理框架。
- 在20个动词、39个物体上实现与现有方法相当的准确率。
- 适合需要透明决策的助人机器人场景,支持人工调试与优化。
在非结构化环境中运行的辅助机器人不仅需识别物体,还需理解其可用性。这要求将语言动作指令与既具备所需功能又可物理抓取的物体关联起来。现有方法多依赖黑箱模型或固定功能标签,限制了可解释性、可控性与可靠性。我们提出CRAFT-E,一种模块化神经符号框架,通过视觉-语言对齐和基于能量的抓握推理,构建包含动词-属性-物体的结构化知识图谱。系统生成可解释的推理路径,揭示影响物体选择的因素,并将抓握可行性作为功能推断的核心部分。我们还构建了一个基准数据集,统一标注动词-物体兼容性、分割结果与抓握候选。在真实机器人上部署全链路流程后,CRAFT-E在静态场景、ImageNet功能检索及包含20个动词和39个物体的真实世界测试中表现优异。该框架在感知噪声下仍具鲁棒性,提供组件级可诊断输出。通过结合符号推理与具身感知,CRAFT-E为功能导向的物体选择提供了可解释且可定制的替代方案,支持助人机器人中的可信决策。
原文摘要 · Abstract (English)
Assistive robots operating in unstructured environments must understand not only what objects are, but what they can be used for. This requires grounding language-based action queries to objects that both afford the requested function and can be physically retrieved. Existing approaches often rely on black-box models or fixed affordance labels, limiting transparency, controllability, and reliability for human-facing applications. We introduce CRAFT-E, a modular neuro-symbolic framework that composes a structured verb-property-object knowledge graph with visual-language alignment and energy-based grasp reasoning. The system generates interpretable grounding paths that expose the factors influencing object selection and incorporates grasp feasibility as an integral part of affordance inference. We further construct a benchmark dataset with unified annotations for verb-object compatibility, segmentation, and grasp candidates, and deploy the full pipeline on a physical robot. CRAFT-E achieves competitive performance in static scenes, ImageNet-based functional retrieval, and real-world trials involving 20 verbs and 39 objects. The framework remains robust under perceptual noise and provides transparent, component-level diagnostics. By coupling symbolic reasoning with embodied perception, CRAFT-E offers an interpretable and customizable alternative to end-to-end models for affordance-grounded object selection, supporting trustworthy decision-making in assistive robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。