arXiv:2608.05490cs.AIcs.LG2026-08

为自主分析代理的错误定位提供可信赖的审计方法,解决误报率与误差溯源难题。

Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability

  • 基于操作序列的异常程度评分,实现错误源头定位
  • 单次错误最多引发一个误报,避免误报扩散
  • 适用于模型不完美或分析选择有偏的场景,理论保障强

自主代理如今可独立完成完整数据分析,包括队列选择、表连接和模型拟合,而无需逐步监督。当分析出错时,需确定是哪个操作导致的。现有方法无需标注错误样本,仅通过学习正确的分析模式来识别偏离行为;但其可靠性尚未被研究。本文对此展开分析:评分方式决定能否定位错误。若按前一操作的意外程度评分,则继承早期错误的操作无法与正确操作区分,导致一次错误仅触发一次警告;若以更长的预期分析路径为参照,单一错误会被分散到多个操作上。本文量化了错误传播范围,并提出在渐进式错误积累下选择合适比较长度的方法。进一步给出控制单个分析中误报比例的程序,只需假设正确分析可交换,无需模型准确。同时量化了模型偏差或分析内容相关选择对保证强度的影响。最后证明任何审计存在根本限制:低于一定幅度的错误无法被归因,因其与正常分析间的波动无异。该极限随收集的正确分析数量增加而缓慢下降,在当前表示维度下,数据量提升百倍也仅使极限改善不足2%,故表示维度是主要约束而非训练数据量。

原文摘要 · Abstract (English)

Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation caused it. A recent approach does this without any labelled mistakes, learning instead from analyses known to be sound and flagging operations that depart from what that model predicts; how reliable such audits are has not been studied. This paper supplies that analysis. The choice of score determines whether an error can be localized at all. If each operation is scored by how surprising it is given the operation immediately preceding it, then operations that merely inherit an earlier error are indistinguishable from correct ones, so one mistake produces one flag; scores computed against a longer reconstruction of the intended analysis instead spread a single mistake across many operations. We quantify how far they spread, and how to choose the comparison length when an error accumulates gradually rather than at once. We then give procedures that control the proportion of falsely flagged operations within a single audited analysis, requiring only that sound analyses be exchangeable rather than that the fitted model be correct, and we quantify how much the guarantees weaken when the model is imperfect or when the analysis was selected for review in a way that depends on its content. Finally we establish a limit on what any such audit can report: errors below a certain magnitude cannot be attributed at all, being indistinguishable from ordinary variation among sound analyses. This limit falls so slowly as more sound analyses are collected that at the representation sizes now in use a hundredfold increase reduces it by under two percent, so the dimension of the representation rather than the volume of training data is the binding constraint.

自主代理错误审计可解释性可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。