arXiv:2510.15236cs.AIcs.CY2025-10

提出用稳定性与核心性评估通用智能,打破传统测试的片面与脆弱。

From Checklists to Clusters: A Homeostatic Account of AGI Evaluation

  • 将通用智能视为抗扰能力集群,以因果核心性权重替代均等评分
  • 设计稳定指数族,验证能力是否持久且可纠错,避免临时表现骗分
  • 无需模型结构即可部署,适合实验室黑箱测试

当前AGI评估虽覆盖多领域,但普遍采用对称权重和快照分数,导致两个问题:(i) 等权处理忽视人类智能研究中各领域的重要性差异;(ii) 快照测试无法区分持久能力与受延迟或压力影响即崩溃的脆弱表现。本文主张通用智能(无论人类或机器)应理解为一种稳态属性集群——一组能力及其维持这些能力在扰动下共存的机制。据此,评估应依据领域对集群稳定的因果核心性加权,并要求跨会话持久性证据。提出两项兼容现有评测的扩展:基于CHC推导权重的中心性优先得分,结合透明敏感性分析;以及一套集群稳定性指数家族,分离出表现持久性、持久学习与错误纠正能力。这些改进在保持多领域广度的同时,降低脆弱性与策略性作弊风险。最后给出可验证预测及无需架构访问的黑箱实验协议。

原文摘要 · Abstract (English)

Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot scores. This creates two problems: (i) equal weighting treats all domains as equally important when human intelligence research suggests otherwise, and (ii) snapshot testing can't distinguish durable capabilities from brittle performances that collapse under delay or stress. I argue that general intelligence -- in humans and potentially in machines -- is better understood as a homeostatic property cluster: a set of abilities plus the mechanisms that keep those abilities co-present under perturbation. On this view, AGI evaluation should weight domains by their causal centrality (their contribution to cluster stability) and require evidence of persistence across sessions. I propose two battery-compatible extensions: a centrality-prior score that imports CHC-derived weights with transparent sensitivity analysis, and a Cluster Stability Index family that separates profile persistence, durable learning, and error correction. These additions preserve multidomain breadth while reducing brittleness and gaming. I close with testable predictions and black-box protocols labs can adopt without architectural access.

AGI评估稳态机制多领域测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。