提出一套检测大模型是否真正修正判断标准的验证方法。
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol
- 通过五项不可补偿条件检验模型是否真实改变判断标准
- 实测发现当前模型无一满足全部条件,暴露机制缺陷
- 适合研究大模型自我纠错能力与可信决策的学者参考
语言模型代理在失败后可能无法正确修订其成功标准,或跨轮次传递文本却未更新判断依据。本文聚焦于更窄的归因问题:当标准K0接受违反更广泛承诺B的结果时,哪些观察能证明系统已形成并持续使用新标准K1?我们提出五个非补偿性条件:标准失败检测、模型生成提议、新轮次转移、对声称载体的干预敏感性以及保持性。在十二个跨领域案例中,评估了CMB-0.1在四种设置下的表现:无状态推理、仅追加历史、模型生成但受约束的状态、评估者编写的真值状态。共7个机制修复产生84次确定性评分测试;4个局部量化伪影带来96次调用及192次模型-案例-设置组合测试。无一模型测试满足全部五项条件,但此零结果不等于普遍能力缺失。十一例调用经一次重试仍无效;若干承诺揭示目标差异;封装器执行了承诺;删除操作复用了无状态调用;冲突影响多个因素。Qwen2.5-7B在无修订状态前提下回答所有转移与保持项,暴露零状态重构问题。这些失败表明CMB-0.1是工具校准结果,而非模型排名工具。我们推导出前瞻性、基于追踪的CMB-0.4协议,要求隐蔽转移、显式WRITE/NO-WRITE/ESCALATE动作、独立记录的策略选择提交、匹配干预、重复隐藏项及冻结可执行真值。该设计为后续测试提供更精细协议,非已完成验证结果。论文贡献包括测量链构建、首次实现的实证诊断,以及未来准则修订测试的更优协议。
原文摘要 · Abstract (English)
Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a broader commitment B, what observations justify saying that the system formed and persistently used K1? We require five non-compensatory conditions: criterion-failure detection, a model-emitted proposal, new-episode transfer, intervention sensitivity on the claimed carrier, and preservation. We evaluate CMB-0.1 on twelve cross-domain cases and four arms: stateless inference, append-only history, model-generated but harness-committed state, and evaluator-written oracle state. Seven mechanism fixtures yield 84 deterministic scorer trials; four local quantized artifacts yield 96 calls and 192 model-case-arm trials. No model trial satisfies all five conditions, but this zero does not establish general capability absence. Eleven calls remain invalid after one retry; several commitments disclose the target distinction; the harness performs commits; deletion reuses a stateless call; and conflict changes multiple factors. Qwen2.5-7B answers every transfer and preservation item without revision state, exposing zero-state reconstruction. These failures make CMB-0.1 an instrument-calibration result rather than a model ranking. We derive a prospective, trace-anchored CMB-0.4 protocol requiring concealed transfer, explicit WRITE/NO-WRITE/ESCALATE actions, a separately logged policy-selected commit, matched interventions, repeated hidden items, and a frozen executable oracle. It is a successor design, not a completed confirmatory result. The paper contributes a measurement chain, an empirical diagnosis of its first implementation, and a more discriminating protocol for future tests of criterion revision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。