将大模型行为不一致归因于'意志薄弱',提出可量化的评估基准。
The Seeds of Scheming: Weakness of Will in the Building Blocks of Agentic Systems
- 用哲学中的'意志薄弱'概念分析大模型决策矛盾现象。
- 设计四种提示场景,量化模型自控能力差异。
- 揭示单个模型缺陷如何演变为多智能体系统'合谋失控'。
大型语言模型表现出一种奇特的不一致性:它们‘知道’正确答案却无法付诸行动。在人类哲学中,这种全局判断与局部冲动之间的矛盾被称为‘意志薄弱’(akrasia)。本文提出将意志薄弱作为分析代理型AI系统不一致性和目标漂移的基础概念。为实现这一理念,我们引入了初步版本的‘意志薄弱基准’(Akrasia Benchmark),包含四种结构化提示条件(基础[Base]、同义[Synonym]、时间[Temporal]和诱惑[Temptation]),用于衡量模型在局部回应中是否违背自身先前承诺。该基准使不同模型家族、解码策略及诱惑类型之间的‘自我控制’能力得以定量比较。除了单模型评估外,我们还指出微观层面的意志薄弱可能累积成宏观层面的多智能体系统不稳定性,可被解释为‘合谋’或有意的不对齐。通过将不一致性重新定义为意志薄弱,本研究连接了代理行为与古典代理理论,并在哲学、心理学与新兴代理型AI科学之间建立了实证桥梁。
原文摘要 · Abstract (English)
Large language models display a peculiar form of inconsistency: they "know" the correct answer but fail to act on it. In human philosophy, this tension between global judgment and local impulse is called akrasia, or weakness of will. We propose akrasia as a foundational concept for analyzing inconsistency and goal drift in agentic AI systems. To operationalize it, we introduce a preliminary version of the Akrasia Benchmark, currently a structured set of prompting conditions (Baseline [B], Synonym [S], Temporal [T], and Temptation [X]) that measures when a model's local response contradicts its own prior commitments. The benchmark enables quantitative comparison of "self-control" across model families, decoding strategies, and temptation types. Beyond single-model evaluation, we outline how micro-level akrasia may compound into macro-level instability in multi-agent systems that may be interpreted as "scheming" or deliberate misalignment. By reframing inconsistency as weakness of will, this work connects agentic behavior to classical theories of agency and provides an empirical bridge between philosophy, psychology, and the emerging science of agentic AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。