arXiv:2512.03072cs.AIcs.LO2025-12

用可追溯的逻辑原子构建可解释、对齐人类价值的通用智能架构

Beyond the Black Box: A Cognitive Architecture for Explainable and Aligned AI

  • 将认知分解为逻辑原子与指认、比较两种基本操作
  • 通过权重=收益×概率实现决策透明,所有值可追溯至初始权重
  • 适合追求可信、可解释AI的研究者与工程团队

当前人工智能范式在可解释性与价值对齐方面面临根本挑战。本文提出一种基于第一性原理的新型认知架构「Weight-Calculatism」,并验证其作为通往通用人工智能(AGI)的可行路径潜力。该架构将认知拆解为不可分割的逻辑原子及两种基本操作:指认与比较。决策过程由可解释的权重计算模型(权重 = 收益 × 概率)形式化,所有数值均可追溯至一组可审计的初始权重。这种原子化分解实现了极致可解释性、对新情境的内在泛化能力以及可追踪的价值对齐。我们通过基于图算法的计算引擎和全局工作区流程实现该架构,并提供初步代码实现与场景验证。结果表明,该架构在前所未见的情境中实现了透明且类人推理与稳健学习,为构建可信、对齐的AGI奠定了实践与理论基础。

原文摘要 · Abstract (English)

Current AI paradigms, as "architects of experience," face fundamental challenges in explainability and value alignment. This paper introduces "Weight-Calculatism," a novel cognitive architecture grounded in first principles, and demonstrates its potential as a viable pathway toward Artificial General Intelligence (AGI). The architecture deconstructs cognition into indivisible Logical Atoms and two fundamental operations: Pointing and Comparison. Decision-making is formalized through an interpretable Weight-Calculation model (Weight = Benefit * Probability), where all values are traceable to an auditable set of Initial Weights. This atomic decomposition enables radical explainability, intrinsic generality for novel situations, and traceable value alignment. We detail its implementation via a graph-algorithm-based computational engine and a global workspace workflow, supported by a preliminary code implementation and scenario validation. Results indicate that the architecture achieves transparent, human-like reasoning and robust learning in unprecedented scenarios, establishing a practical and theoretical foundation for building trustworthy and aligned AGI.

可解释AI认知架构价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。