arXiv:2509.23757cs.AIcs.CV2025-09

用多智能体协作让视觉推理过程透明可懂。

Transparent Visual Reasoning via Object-Centric Agent Collaboration

  • 基于物体中心表示与博弈论推理,实现可解释的决策过程。
  • 在两个多物体数据集上表现媲美顶尖黑盒模型。
  • 用户研究证实其解释更直观可信,适合医疗等高风险场景。

可解释人工智能的核心挑战之一是生成基于人类可理解概念的解释。为此,我们提出OCEAN(基于智能体协商的物体中心解释框架),一个基于物体中心表征和透明多智能体推理过程的全新可解释框架。博弈论推理促使智能体就一致且有区分性的证据达成共识,从而实现忠实且可解释的决策过程。我们在两个诊断型多物体数据集上端到端训练OCEAN,并与标准视觉分类器及流行的后验解释工具(如GradCAM、LIME)进行对比。结果表明,OCEAN在性能上媲美最先进黑盒模型,且推理过程真实可信。用户研究表明,参与者普遍认为OCEAN的解释更直观、更值得信赖。

原文摘要 · Abstract (English)

A central challenge in explainable AI, particularly in the visual domain, is producing explanations grounded in human-understandable concepts. To tackle this, we introduce OCEAN (Object-Centric Explananda via Agent Negotiation), a novel, inherently interpretable framework built on object-centric representations and a transparent multi-agent reasoning process. The game-theoretic reasoning process drives agents to agree on coherent and discriminative evidence, resulting in a faithful and interpretable decision-making process. We train OCEAN end-to-end and benchmark it against standard visual classifiers and popular posthoc explanation tools like GradCAM and LIME across two diagnostic multi-object datasets. Our results demonstrate competitive performance with respect to state-of-the-art black-box models with a faithful reasoning process, which was reflected by our user study, where participants consistently rated OCEAN's explanations as more intuitive and trustworthy.

可解释AI多智能体视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。