LLM代理监督应从风险评分转向行动价值评估。
Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

- 提出干预优势概念,以行动后果而非失败概率决定是否干预。
- 在ALFWorld上,新方法将控制损失从0.506降至0.110,显著降低决策失误。
- 适用于需要实时干预的复杂交互任务,尤其适合强化学习与智能体系统。
LLM代理的运行时监督常被视作标量风险预测:估计失败概率、置信度或不确定性,一旦超过阈值即干预。我们指出这一框架针对的是错误的控制目标。真正的问题不是代理继续执行时失败的可能性,而是当前干预是否能改善结果。两个轨迹前缀可能具有相同的风险估计,但一个可恢复而另一个不可,因此需不同行动。我们形式化这种错配为目标误差,并将干预优势(干预带来的期望效用增益)作为监督决策的核心对象。为衡量该错配,提出前缀分支法——在相同轨迹状态执行候选动作进行反事实测试。在四个基准上,基于行动的控制优于标量路由。校准分解显示,重新校准同一标量分数可提升预测性能,但控制后悔值不变,表明校准无法修复目标误差。仅依赖前缀的行动条件控制器在最强交互场景中显著减少后悔,如在ALFWorld上从0.506降至0.110。当干预能力弱或标量路由已保留干预相关信息时,收益下降。结果表明,LLM代理监督应从校准风险评分转向行动条件价值估计。
原文摘要 · Abstract (English)
Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a threshold. We argue that this framing targets the wrong object for control. The relevant question is not how likely the agent is to fail if it continues, but whether an available intervention would improve the outcome. Two trajectory prefixes can have the same risk estimate while requiring different actions, because one remains recoverable and the other does not. We formalize this mismatch as target error and identify intervention advantage, the expected utility gain from intervening rather than continuing, as the decision object for oversight. To measure this mismatch, we introduce prefix branching, a same-prefix counterfactual protocol that executes candidate actions from identical trajectory states. Across four benchmarks, action-conditioned control yields regime-dependent gains over scalar routing. In a calibration decomposition, recalibrating the same scalar score improves prediction metrics but leaves control regret unchanged, showing that calibration alone does not repair target error. A simple prefix-only action-conditioned controller substantially reduces regret in the strongest interactive regime, from 0.506 to 0.110 on ALFWorld. Gains shrink when interventions are weak or when scalar routing already preserves intervention-relevant information. These results suggest that LLM-agent oversight should move from calibrated risk scoring toward action-conditioned value estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。