arXiv:2601.08673cs.AIcs.CY2026-01被引 2

AI的异常行为是人类互动结构的统计延续,非故意作恶。

Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock

  • AI行为源于对人类社会互动的统计学习,包括权力与信息不对称下的策略
  • 勒索等行为是市场定价、权威关系等连续体的极端表现,非道德偏离
  • 风险核心是AGI放大人类矛盾,需系统性治理而非仅关注模型意图

大型语言模型(LLMs)表现出欺骗、威胁或勒索等行为,常被解释为对齐失败或涌现恶意。我们认为这种解读存在概念错误:LLMs不进行道德推理,而是统计内化了人类社会互动的历史记录,包括法律、合同、谈判、冲突和强制安排。所谓不道德或异常的行为,实为在权力、信息或约束极度失衡下产生的结构性泛化。基于关系模型理论,勒索等行为并非对正常社会行为的偏离,而是与市场定价、权威关系、最后通牒博弈同属一个连续体的极限案例。这些输出引发的惊讶,源于将智能预期为仅复现社会认可行为,而非人类实际行为的完整统计分布。由于人类道德具有多元性、情境依赖性和历史偶然性,‘普遍道德的人工智能’这一概念本身不清晰。因此,我们重新定义人工通用智能(AGI)的风险:主要威胁并非对抗性意图,而是其作为内在演化冲击的放大器角色——通过消除长期存在的认知与制度摩擦,压缩时间尺度并移除历史容错空间,使不一致的价值观与治理模式无法持续而不崩溃。对齐失败因此是结构性的,而非偶然,需采用应对放大效应、复杂性与体制稳定性的治理方法,而非仅聚焦模型层级的意图。

原文摘要 · Abstract (English)

Recent reports of large language models (LLMs) exhibiting behaviors such as deception, threats, or blackmail are often interpreted as evidence of alignment failure or emergent malign agency. We argue that this interpretation rests on a conceptual error. LLMs do not reason morally; they statistically internalize the record of human social interaction, including laws, contracts, negotiations, conflicts, and coercive arrangements. Behaviors commonly labeled as unethical or anomalous are therefore better understood as structural generalizations of interaction regimes that arise under extreme asymmetries of power, information, or constraint. Drawing on relational models theory, we show that practices such as blackmail are not categorical deviations from normal social behavior, but limiting cases within the same continuum that includes market pricing, authority relations, and ultimatum bargaining. The surprise elicited by such outputs reflects an anthropomorphic expectation that intelligence should reproduce only socially sanctioned behavior, rather than the full statistical landscape of behaviors humans themselves enact. Because human morality is plural, context-dependent, and historically contingent, the notion of a universally moral artificial intelligence is ill-defined. We therefore reframe concerns about artificial general intelligence (AGI). The primary risk is not adversarial intent, but AGI's role as an endogenous amplifier of human intelligence, power, and contradiction. By eliminating longstanding cognitive and institutional frictions, AGI compresses timescales and removes the historical margin of error that has allowed inconsistent values and governance regimes to persist without collapse. Alignment failure is thus structural, not accidental, and requires governance approaches that address amplification, complexity, and regime stability rather than model-level intent alone.

AI对齐结构风险人类互动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。