arXiv:2507.14426cs.CV2025-07中稿 · NeSy 2025被引 1

让机器理解物体能否执行某个动作,且过程可解释。

CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding

  • 融合常识知识与视觉特征,用能量机制迭代优化判断
  • 在无标签多物体场景中准确识别可执行动作的物体
  • 适合需要透明决策的智能系统开发

我们提出CRAFT,一种可解释的神经符号框架,用于视觉功能性使用关系的定位,旨在识别场景中能实现特定动作(如“切”)的物体。CRAFT结合来自ConceptNet的结构化常识先验与语言模型,以及来自CLIP的视觉证据,通过基于能量的推理循环,迭代优化预测结果。该过程生成透明、目标驱动的决策,实现符号结构与感知结构的对齐。在多物体、无标签设置下的实验表明,CRAFT在提升准确率的同时增强了可解释性,为构建鲁棒且可信的场景理解系统迈出关键一步。

原文摘要 · Abstract (English)

We introduce CRAFT, a neuro-symbolic framework for interpretable affordance grounding, which identifies the objects in a scene that enable a given action (e.g., "cut"). CRAFT integrates structured commonsense priors from ConceptNet and language models with visual evidence from CLIP, using an energy-based reasoning loop to refine predictions iteratively. This process yields transparent, goal-driven decisions to ground symbolic and perceptual structures. Experiments in multi-object, label-free settings demonstrate that CRAFT enhances accuracy while improving interpretability, providing a step toward robust and trustworthy scene understanding.

可解释性视觉理解神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。