arXiv:2603.29654cs.LGcs.AI2026-03

发现机器学习中的概念矛盾,帮助对齐人类与机器的理解。

Concept frustration: Aligning human concepts and machine representations

  • 用几何方法对比人类概念与模型内部表示的不一致
  • 实验证明基础模型中存在可检测的概念矛盾现象
  • 适合关注可解释AI安全与推理对齐的研究者

将人类可理解的概念与现代机器学习系统内部表征对齐,仍是可解释AI的核心挑战。本文提出一种几何框架,用于比较监督式人类概念与基础模型嵌入中提取的无监督中间表示。受科学发现中概念跃迁的启发,我们形式化了“概念挫折”:当一个未观测概念引发已知概念间关系,而这些关系在现有本体论中无法自洽时产生的矛盾。我们开发了任务对齐的相似性度量,可在任务对齐几何中检测概念挫折,而传统欧氏比较则失效。在线性高斯生成模型下,我们推导出贝叶斯最优概念分类器的闭式表达,将预测信号分解为已知-已知、已知-未知和未知-未知三部分,并解析识别出挫折影响性能的位置。在合成数据及真实语言与视觉任务上的实验表明,概念挫折可在基础模型表示中被检测到,且引入挫折概念可重构可解释模型中概念表征的几何结构,更好地对齐人类与机器推理。结果揭示了一种诊断不完备概念本体的原理性框架,对高风险应用中安全可解释AI的发展与验证具有意义。

原文摘要 · Abstract (English)

Aligning human-interpretable concepts with the internal representations learned by modern machine learning systems remains a central challenge for interpretable AI. We introduce a geometric framework for comparing supervised human concepts with unsupervised intermediate representations extracted from foundation model embeddings. Motivated by the role of conceptual leaps in scientific discovery, we formalise the notion of concept frustration: a contradiction that arises when an unobserved concept induces relationships between known concepts that cannot be made consistent within an existing ontology. We develop task-aligned similarity measures that detect concept frustration between supervised concept-based models and unsupervised representations derived from foundation models, and show that the phenomenon is detectable in task-aligned geometry while conventional Euclidean comparisons fail. Under a linear-Gaussian generative model we derive a closed-form expression for Bayes-optimal concept-based classifier accuracy, decomposing predictive signal into known-known, known-unknown and unknown-unknown contributions and identifying analytically where frustration affects performance. Experiments on synthetic data and real language and vision tasks demonstrate that frustration can be detected in foundation model representations and that incorporating a frustrating concept into an interpretable model reorganises the geometry of learned concept representations, to better align human and machine reasoning. These results suggest a principled framework for diagnosing incomplete concept ontologies and aligning human and machine conceptual reasoning, with implications for the development and validation of safe interpretable AI for high-risk applications.

可解释AI概念对齐基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。