arXiv:2604.08588cs.LGcs.AI2026-04

让大模型学会何时自己处理、何时上报,提升自动化可靠性。

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models

论文配图:Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models
图 1 · 摘自论文原文
  • 用不确定性决策框架评估模型该自己做还是上报
  • 不同模型的上报阈值差异大,且自我评估常不准
  • 通过思维链微调可让模型稳定遵循正确上报策略

有效的自动化依赖于判断何时自主处理、何时上报。本文将此建模为不确定性下的决策:大模型生成预测并估算其正确概率,再比较行动与上报的预期成本。在需求预测、内容推荐、内容审核、贷款审批和自动驾驶五个真实人类决策场景中,测试多个模型家族,发现模型隐含的上报阈值存在显著差异,且不受架构或规模影响,自我置信度也呈现模型特异性偏差。进一步测试了调整成本比、提供准确率信号、以及训练模型遵循特定上报规则等干预手段。提示工程对推理类模型效果有限,而基于思维链的监督微调(SFT)产生最稳健的策略,能在不同数据集、成本比例、提示格式及未见领域间泛化。结果表明,上报行为是模型特有属性,部署前需评估;强化模型对不确定性和决策成本的显式推理,有助于实现鲁棒对齐。

原文摘要 · Abstract (English)

Effective automation hinges on deciding when to act and when to escalate. We model this as a decision under uncertainty: an LLM forms a prediction, estimates its probability of being correct, and compares the expected costs of acting and escalating. Using this framework across five domains of recorded human decisions-demand forecasting, content recommendation, content moderation, loan approval, and autonomous driving-and across multiple model families, we find marked differences in the implicit thresholds models use to trade off these costs. These thresholds vary substantially and are not predicted by architecture or scale, while self-estimates are miscalibrated in model-specific ways. We then test interventions that target this decision process by varying cost ratios, providing accuracy signals, and training models to follow the desired escalation rule. Prompting helps mainly for reasoning models. SFT on chain-of-thought targets yields the most robust policies, which generalize across datasets, cost ratios, prompt framings, and held-out domains. These results suggest that escalation behavior is a model-specific property that should be characterized before deployment, and that robust alignment benefits from training models to reason explicitly about uncertainty and decision costs.

大模型自动化决策对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。