arXiv:2605.24727cs.AIcs.CL2026-05

AI解释存在根本性局限,无法同时满足环境复杂、性能好、可解释和完全忠实四条件。

Fundamental Limitation in Explaining AI

  • 提出四重困境理论:四个理想解释条件不可兼得。
  • 证明在多数实际场景下,必须放弃解释的完全忠实性。
  • 为AI治理提供新思路:解释必然不完整,应聚焦关键部分。

尽管大模型如LLMs和扩散模型已取得实际成功,公众机构日益重视AI可解释性。然而,现有解释方法并未设计用于提供大规模AI系统行为的完全忠实解释。虽然完全忠实且可解释的解释对AI治理可能有益,但其理论可行性尚不明确。本文数学证明了AI解释的根本性四重困境:一个AI及其解释无法同时满足以下四个条件——1)操作环境的复杂性;2)AI性能的良好表现;3)解释的可解释性;4)解释的完全忠实性。该困境表明,在无法改变环境、不牺牲良好性能或可解释性的多数应用场景中,必须放弃解释的完全忠实性,转而专注于解释对应用至关重要的部分。因此,该四重困境意味着AI治理应基于解释始终不完全忠实的前提进行设计。

原文摘要 · Abstract (English)

While large-scale models such as LLMs and diffusion models have achieved practical success, public institutions have emphasized the importance of explainability in AI. Existing methods for explaining AI, however, are not designed to provide completely faithful explanations of the behavior of large-scale AI systems. Although a completely faithful and interpretable explanation of the behavior of an AI system might be useful for AI governance, it has not been known whether providing such an explanation is theoretically possible. In this paper, we mathematically prove a fundamental quadrilemma in explaining AI, stating that AI and its explanation cannot satisfy the following four conditions simultaneously: 1) the complexity of the operation environment, 2) the goodness of the AI's performance, 3) the interpretability of the AI's explanation, and 4) the complete faithfulness of the AI's explanation. This quadrilemma suggests that, in most applications where we cannot change the environment or sacrifice good AI performance and an interpretable explanation, we should give up complete faithfulness of explanations and should instead aim to explain only the parts that are important for applications. As a consequence, the quadrilemma implies that AI governance should be designed on the premise that the faithfulness of AI explanations is always incomplete.

AI解释理论局限可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。