arXiv:2608.16003cs.AIcs.CL2026-08

修复过程会改变大模型验证器的判断标准,使其更宽容。

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

  • 用修复后的上下文影响验证模型,降低误报率。
  • 15组测试中误报率下降2.8至11.5个百分点,降幅9%~25%。
  • 适合关注模型安全性和自动化检测可信度的研究者。

自动化检查流水线中,一个语言模型担任检查者,另一个(或同一模型)负责修复。我们探究这种配置是否改变检查者的行为。在保持任务字节完全一致的前提下,测量人类验证正确的ProcessBench样本上的误报率,发现已包含完整审计-修复流程的上下文可使15组模型与表述组合中的误报率全部降低2.8至11.5个百分点,相比长度匹配的无审计对照组减少9%至25%。该趋势与累积信息文献预测相反:当审计报告存在错误时,误报率进一步下降,在五种表述下均成立,尽管负面偏差理论预测应增加标记。分解分析显示,修复内容和审计结论具有互补作用,不同组件对不同模型家族产生影响。信号检测分析表明,变化源于阈值移动而非辨别力提升——15组中有15组阈值改变且修正后仍维持,而辨别力指标d'仅在13组中存活,但其测试本身灵敏度仅为一半。人工审查50个误报案例发现82%为明显错误,因此在此操作点上阈值变化未必有害。开启推理模式后,效应相对大小在两个测试模型上保持不变,阈值也依然有效。

原文摘要 · Abstract (English)

Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.

大模型验证误报控制上下文影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。