arXiv:2512.19155cs.AI2025-12被引 1

用人工模型测试意识理论,发现三类理论可互补而非互斥。

Can We Test Consciousness Theories on AI? Ablations, Markers, and Robustness

  • 构建人工代理模拟不同意识理论,通过精确删减测试功能影响。
  • 各理论在任务表现与元认知上出现分离,支持分层功能假说。
  • 适合研究意识机制的计算神经科学与AI伦理学者参考。

寻找可靠的意识指标已分化为多个理论阵营(全局工作空间理论GWT、整合信息理论IIT、高阶理论HOT),各自提出不同的神经标志。本文采用合成神经现象学方法:构建体现这些机制的人工智能代理,通过生物系统无法实现的精确架构删减,检验其功能后果。三个实验表明,这些理论描述的是互补的功能层级而非竞争关系。实验1中,去除自我模型导致元认知校准丧失,但一阶任务性能保留,形成类似人类盲视的合成现象,符合HOT预测。实验2中,工作空间容量对信息可及性具有因果必要性:完全删除工作空间导致相关标志质变崩溃,部分减少则呈渐进退化,符合GWT点火框架。实验3揭示广播放大效应:GWT式广播会放大内部噪声,造成极端脆弱性;而B2系列代理在同一扰动下保持鲁棒,该特性在关闭自我模型/仅读取工作空间的控制条件下仍存,提示该鲁棒性非仅源于$z_{ ext{self}}$压缩。我们还报告一个明确的负面结果:在工作空间瓶颈下,原始扰动复杂性(PCI-A)下降,警示不可简单将IIT相关代理迁移至工程化智能体。结果暗示一种分层设计原则:GWT提供广播能力,HOT提供质量控制。强调这些代理本身并不具备意识,而是用于验证意识理论功能预测的参照实现。

原文摘要 · Abstract (English)

The search for reliable indicators of consciousness has fragmented into competing theoretical camps (Global Workspace Theory (GWT), Integrated Information Theory (IIT), and Higher-Order Theories (HOT)), each proposing distinct neural signatures. We adopt a synthetic neuro-phenomenology approach: constructing artificial agents that embody these mechanisms to test their functional consequences through precise architectural ablations impossible in biological systems. Across three experiments, we report dissociations suggesting these theories describe complementary functional layers rather than competing accounts. In Experiment 1, a no-rewire Self-Model lesion abolishes metacognitive calibration while preserving first-order task performance, yielding a synthetic blindsight analogue consistent with HOT predictions. In Experiment 2, workspace capacity proves causally necessary for information access: a complete workspace lesion produces qualitative collapse in access-related markers, while partial reductions show graded degradation, consistent with GWT's ignition framework. In Experiment 3, we uncover a broadcast-amplification effect: GWT-style broadcasting amplifies internal noise, creating extreme fragility. The B2 agent family is robust to the same latent perturbation; this robustness persists in a Self-Model-off / workspace-read control, cautioning against attributing the effect solely to $z_{\text{self}}$ compression. We also report an explicit negative result: raw perturbational complexity (PCI-A) decreases under the workspace bottleneck, cautioning against naive transfer of IIT-adjacent proxies to engineered agents. These results suggest a hierarchical design principle: GWT provides broadcast capacity, while HOT provides quality control. We emphasize that our agents are not conscious; they are reference implementations for testing functional predictions of consciousness theories.

意识理论人工智能神经机制实验验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。