arXiv:2604.18245cs.LG2026-04

提出纠正与污染率新视角,量化大模型调用中错误流动。

Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols

论文配图:Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols
图 1 · 摘自论文原文
  • 通过配对审计法记录操作前后成功状态,区分纠正失败与污染成功。
  • 高基线准确率下需更高纠正率才能净增收益,否则反而有害。
  • 适用于评估模型调用组合、输入选择及校准策略的有效性。

大型语言模型在多调用协议中运行,但新增调用通常仅以整体效果评估,无法区分纠正失败输出与污染初始成功输出的区别。本文提出一种配对审计方法,在同一任务上基于二元规则记录操作前后的成功状态。纠正率与污染率精确刻画净变化:收益来自被纠正的失败,损失源于被污染的成功。当基线成功率上升时,达到盈亏平衡所需的纠正率急剧提高,因此相同行为可能提升中等基线群体表现,却损害高基线群体。应用于已发表的GPT-4 GSM8K结果,该框架可从聚合准确率推断样本内纠正与污染次数。研究进一步探讨校准估计是否能预测未观测结果,操作所获信息量如何影响率值,以及连续测量是否可叠加。校准估计在独立样本上跟踪生成器的准确性;在按生成器定义分组重加权时,全局估计可能失效,而组内估计可降低平均预测误差。在GSM8K上,可观测特征在当前样本量下仅捕捉有限变异,故组级应用规则仍具探索性。重新排列固定四候选集会同时改变两率,准确率变化取决于是否存在正确替代项;在MBPP上,增加辅助工具使357个初始通过程序的污染数从28升至100。样本内直接与复合转移估计平均接近;以初始成功行的支撑权重进行加权,显著减少保留样本差异。纠正与污染是针对特定操作的度量,非模型常量,可指导输入选择、应用决策与组合测试。

原文摘要 · Abstract (English)

Large language models operate in protocols containing multiple calls, yet added calls are usually evaluated only by their net effect. That summary cannot distinguish correcting unsuccessful outputs from corrupting initially successful ones. We develop a paired audit recording success before and after a specified operation on the same tasks under one binary rule. Correction and corruption rates exactly account for the net change: gains come from corrected failures and losses from corrupted successes. The break-even correction requirement rises sharply with baseline success, so the same behavior can improve a moderate-baseline population but harm a high-baseline one. Applied to published GPT-4 GSM8K results, the framework bounds within-sample correction and corruption counts from aggregate accuracies. We ask if calibration estimates predict unobserved outcomes, how rates change with information supplied to an operation, and whether successive measurements combine. Calibration estimates track accuracy on a disjoint sample from the same generator. Under reweighting of generator-defined groups, pooled estimates can fail while group-specific estimates reduce average prediction error. On GSM8K, observable features capture limited variation at these sample sizes, so the group-level application rule remains exploratory. Reordering a fixed four-candidate set changes both rates, with the accuracy effect depending on whether a correct alternative is available; on MBPP, adding a helper artifact raises corruptions among 357 initially passing programs from 28 to 100. Direct and composed transition estimates are close on average within samples; weighting initial-success rows by support substantially reduces held-out discrepancy. Correction and corruption are measurements of specified operations, not model constants, and can guide input choices, application decisions, and composition tests.

大模型评估错误分析调用协议校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。