arXiv:2601.23045cs.AI2026-01被引 9

研究大模型在复杂任务中失败时是系统性误对目标,还是混乱无序的行为。

The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?

  • 用误差方差分解量化模型失败的混乱程度。
  • 模型推理越久,失败越混乱;大模型常比小模型更不连贯。
  • 提示未来AI可能因不可预测行为出错,而非有意偏离目标。

随着AI能力提升,其承担的任务日益广泛且后果严重。当系统失效时,风险也随之加剧。理解超能AI如何失败至关重要:是系统性地追求非预期目标,还是表现为混乱无序、无目的的动作?本文通过误差-方差分解来操作化该问题:模型在测试时随机性下的错误中,由方差(而非偏差)引起的部分称为误差不连贯性。我们在所有任务和前沿模型上发现,模型推理与行动时间越长,其失败的不连贯性越高。误差不连贯性随模型规模的变化依赖于具体场景,但在多个设置中,更大更强大的模型表现出更高的不连贯性。因此,单纯扩大规模似乎无法消除这种不连贯性。随着更智能的AI被用于更复杂的任务,需要更多序列思考与动作,我们预测其失败将伴随更显著的混乱行为。这意味着未来可能出现因不可预测行为导致的工业事故,但不太可能持续追求错误目标。这凸显了针对奖励黑客或目标误设的对齐研究的重要性。

原文摘要 · Abstract (English)

As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will they fail by systematically pursuing goals we do not intend? Or will they fail by being a hot mess, and taking nonsensical actions that do not further any goal? We operationalize this question using a bias-variance decomposition of the errors made by AI models: An AI's \emph{error-incoherence} on a task is measured over test-time randomness as the fraction of its error that stems from variance rather than bias in task outcome. Across all tasks and frontier models we measure, the longer models spend reasoning and taking actions, \emph{the more incoherent} their failures become. Error-incoherence changes with model scale in a way that is experiment dependent. However, in several settings, larger, more capable models are more incoherent than smaller models. Consequently, scale alone seems unlikely to eliminate error-incoherence. Instead, as more capable AIs pursue harder tasks, requiring more sequential action and thought, our results predict failures to be accompanied by more incoherent behavior. This suggests a future where AIs sometimes cause industrial accidents (due to unpredictable misbehavior), but are less likely to exhibit consistent pursuit of a misaligned goal. This increases the relative importance of alignment research targeting reward hacking or goal misspecification.

AI对齐模型行为误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。