arXiv:2504.02058cs.AI2025-04被引 1

揭示了对齐研究中的认知封闭如何阻碍创新,导致不可逆的对齐失效。

Epistemic Closure and the Irreversibility of Misalignment: Modeling Systemic Barriers to Alignment Innovation

  • 构建认知封闭模型,分析评估体系如何过滤掉新范式。
  • 实证发现92%的去中心化集体智能提案被现有系统拒斥且未被审阅。
  • 强调唯有递归建模封闭机制本身才能突破僵局,适合安全研究者参考。

当前人工智能通用智能(AGI)安全发展多依赖基于公理形式主义、可解释性与实证验证的共识对齐方法,但这些方法可能因结构缺陷而无法识别或接纳超出其认知框架的新方案。本文提出一个加权的认知封闭功能模型,整合认知、制度、社会与基础设施过滤机制,使许多对齐提案在现有评估体系中‘不可读’。该模型基于理论与实证双重支持,包括由AI系统执行的元分析,揭示去中心化集体智能(DCI)框架在学术评审中遭遇92%的拒绝率和零实质审阅。我们主张,对类似DCI的模型持续忽视并非偶然疏漏,而是结构性吸引子,复现了我们试图避免的对齐风险。若不采用诸如DCI般具有递归修正能力的模型,人类将不可避免地走向不可逆的对齐失败。本文自身通过模拟评审与正式发表流程的发展,成为支持核心论点的案例:唯有递归建模封闭本身的约束,方能打破封闭。

原文摘要 · Abstract (English)

Efforts to ensure the safe development of artificial general intelligence (AGI) often rely on consensus-based alignment approaches grounded in axiomatic formalism, interpretability, and empirical validation. However, these methods may be structurally unable to recognize or incorporate novel solutions that fall outside their accepted epistemic frameworks. This paper introduces a functional model of epistemic closure, in which cognitive, institutional, social, and infrastructural filters combine to make many alignment proposals illegible to existing evaluation systems. We present a weighted closure model supported by both theoretical and empirical sources, including a meta-analysis performed by an AI system on patterns of rejection and non-engagement with a framework for decentralized collective intelligence (DCI). We argue that the recursive failure to assess models like DCI is not just a sociological oversight but a structural attractor, mirroring the very risks of misalignment we aim to avoid in AGI. Without the adoption of DCI or a similarly recursive model of epistemic correction, we may be on a predictable path toward irreversible misalignment. The development and acceptance of this paper, first through simulated review and then through formal channels, provide a case study supporting its central claim: that epistemic closure can only be overcome by recursive modeling of the constraints that sustain it.

对齐难题认知封闭递归修正安全研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。