arXiv:2507.07532cs.LGcs.AI2025-07被引 1

用概念编码让复杂图像也能被可验证的AI模型理解。

Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings

  • 用最小监督方法提取输入的结构化概念编码,供验证模型使用。
  • 在高维逻辑复杂数据上,性能优于传统概念模型和像素基基线。
  • 适合关注可解释性与形式验证的AI研究者,尤其适用于图像任务。

尽管证明者-验证者游戏(PVGs)为非线性分类模型提供了可验证性的前景,但尚未应用于高维图像等复杂输入。相反,表达性强的概念编码能有效将此类数据转化为可解释概念,但通常仅用于低容量线性预测器。本文提出神经概念验证器(NCV),融合PVG的形式可验证性与概念编码对复杂高维输入的可解释处理能力。NCV利用最近的最小监督概念发现模型从原始输入中提取结构化概念编码,证明者从中选择子集,验证者则仅基于这些编码进行决策,且验证者为非线性预测器。实验表明,NCV在高维、逻辑复杂的数据集上优于经典概念模型和基于像素的PVG分类器基线,有助于缓解捷径行为。整体上,本工作展示了向概念级可验证AI迈进的重要一步。

原文摘要 · Abstract (English)

While Prover-Verifier Games (PVGs) offer a promising path toward verifiability in nonlinear classification models, they have not yet been applied to complex inputs such as high-dimensional images. Conversely, expressive concept encodings effectively allow to translate such data into interpretable concepts but are often utilised in the context of low-capacity linear predictors. In this work, we push towards real-world verifiability by combining the strengths of both approaches. We introduce Neural Concept Verifier (NCV), a unified framework combining PVGs for formal verifiability with concept encodings to handle complex, high-dimensional inputs in an interpretable way. NCV achieves this by utilizing recent minimally supervised concept discovery models to extract structured concept encodings from raw inputs. A prover then selects a subset of these encodings, which a verifier, implemented as a nonlinear predictor, uses exclusively for decision-making. Our evaluations show that NCV outperforms classic concept-based models and pixel-based PVG classifier baselines on high-dimensional, logically complex datasets and helps mitigate shortcut behavior. Overall, we demonstrate NCV as a promising step toward concept-level, verifiable AI.

可验证AI概念编码图像理解形式验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。