让机器理解物体能否执行某个动作,且过程可解释。
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
- 融合常识知识与视觉特征,用能量机制迭代优化判断
- 在无标签多物体场景中准确识别可执行动作的物体
- 适合需要透明决策的智能系统开发
我们提出CRAFT,一种可解释的神经符号框架,用于视觉功能性使用关系的定位,旨在识别场景中能实现特定动作(如“切”)的物体。CRAFT结合来自ConceptNet的结构化常识先验与语言模型,以及来自CLIP的视觉证据,通过基于能量的推理循环,迭代优化预测结果。该过程生成透明、目标驱动的决策,实现符号结构与感知结构的对齐。在多物体、无标签设置下的实验表明,CRAFT在提升准确率的同时增强了可解释性,为构建鲁棒且可信的场景理解系统迈出关键一步。
原文摘要 · Abstract (English)
We introduce CRAFT, a neuro-symbolic framework for interpretable affordance grounding, which identifies the objects in a scene that enable a given action (e.g., "cut"). CRAFT integrates structured commonsense priors from ConceptNet and language models with visual evidence from CLIP, using an energy-based reasoning loop to refine predictions iteratively. This process yields transparent, goal-driven decisions to ground symbolic and perceptual structures. Experiments in multi-object, label-free settings demonstrate that CRAFT enhances accuracy while improving interpretability, providing a step toward robust and trustworthy scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。