arXiv:2509.21654cs.LGcs.AI2025-09被引 2

AI无法同时做到准确、可信且达到人类推理水平。

Limitations on Accurate, Trusted, Human-level Reasoning

  • 定义准确:能拒绝预测时绝不出错;可信:假设系统准确。
  • 证明存在人类易解但AI无法解决的任务,即使系统准确可信。
  • 用哥德尔与图灵思想类比,揭示可信系统的内在局限性。

我们发现,在对准确性、可信性和人类级推理进行严格数学定义的前提下,人工智能系统难以同时实现这三个目标。准确性指系统在有能力拒答时从不做出错误判断;可信指假设系统是准确的;人类级推理指系统始终达到或超过人类能力。核心结论是:一个准确且可信的系统不可能具备人类级推理能力——对于这类系统,存在人类可轻松且可证明解决的任务,但系统无法处理。证明思路借鉴了哥德尔不完备性定理和图灵对停机问题不可判定性的证明,可视为对这些经典结果的现代诠释。关键在于形式化‘可信’概念,从而将系统内在属性(准确性)与其认知状态(被信任)区分开。

原文摘要 · Abstract (English)

We identify a fundamental incompatibility between the goals of accuracy, trust, and human-level reasoning in artificial intelligence (AI) systems, for strict mathematical definitions of these notions. We define accuracy of a system as the property that it never makes any false claims when it has the ability to abstain from making a prediction on any input, and trust as the assumption that the system is accurate. We define human-level reasoning as the property of an AI system always matching or exceeding human capability. Our core finding is that -- for our formal definitions of these notions -- an accurate and trusted AI system cannot be a human-level reasoning system: for such an accurate, trusted system there are task instances which are easily and provably solvable by a human but not by the system. Our proofs draw parallels to Gödel's incompleteness theorems and Turing's proof of the undecidability of the halting problem, and can be regarded as interpretations of Gödel's and Turing's results. Key to our proof is the formalization of the notion of trust, which allows us to separate the intrinsic property of a system (being accurate) from its epistemic status (being trusted).

AI理论可信度逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。