arXiv:2606.22495cs.AI2026-06

提出环境确定性是制约智能体长链任务成功的核心瓶颈。

Grounded Scaling: Why Agentic AI Needs Deterministic Environments

  • 构建确定性-效率边界,揭示任务成功率随步数指数衰减
  • 发现奖励不完美时验证者会劣化,导致飞轮效应受限
  • 提供可量化的确定性成熟度模型,适配各类智能体系统

长链智能体执行在人类容忍型环境中呈指数失败:当每步确定性δ<1时,k步链任务成功率降为δ^k。当前关于AGI到ASI的扩展争论(Genewein等,2026)将进展视为算力增长与数据墙、抽象屏障、具身瓶颈、多智能体信任等摩擦之间的竞赛;我们主张环境确定性是贯穿这四类摩擦的互补关键轴,尤其适用于经济、物理或多方结算可验证的智能体任务。三个形式化结果确立了该范式:链任务成功率的确定性-效率边界、不完美奖励下验证者退化的验证者古德哈特阈值、以及环境侧技能演化的收敛条件。框架被具体化为五个可测属性的供应确定性指数(SCI)、五级确定性成熟度模型(DMM)作为采纳阶梯,及一套可证伪的开放问题程序(OQ1-OQ5),其零结果将迫使观点修正。立场平台无关,回应了仿真到现实充分性、对齐充分性、人工智能即普通技术三种对立观点。

原文摘要 · Abstract (English)

Long-chain agent execution fails exponentially in environments designed for human tolerance: with per-step determinism $δ< 1$, $k$-step chain success degrades as $δ^k$. The AGI-to-ASI scaling debate (Genewein et al., 2026) has so far framed progress as a race between compute growth and a list of frictions (data wall, abstraction barrier, embodied bottleneck, multi-agent trust); we argue that environment determinism is a complementary binding axis cutting across all four, for the broad class of agentic AI tasks whose outcomes are verifiable economically, physically, or through multi-party settlement. Three formal results pin down the regime: a Determinism-Efficiency Bound on chain-task success, a Verifier-Goodharting Floor on flywheel ceilings under imperfect rewards, and a convergence condition for environment-side skill evolution. We operationalise the framework as a Supply Certainty Index (SCI) over five measurable properties, a five-level Determinism Maturity Model (DMM) as adoption ladder, and a falsifiable open-question programme (OQ1-OQ5) with explicit null results that would force retraction. The position is platform-agnostic. We engage three competing positions: sim-to-real sufficiency, alignment sufficiency, and AI-as-normal-technology.

智能体确定性长链推理评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。