用简单检查发现数据科学代理的虚假结论。
Sanity Checks for Agentic Data Science

- 基于预测性-可计算性-稳定性框架设计轻量级验证方法
- 在11个真实数据集上发现6个结论缺乏可靠支持
- 适合评估AI生成分析结果可信度的研究者与从业者
自主式数据科学(ADS)系统能力迅速提升,如OpenAI Codex已能直接分析数据并回答统计问题。然而,这些系统可能得出看似合理实则错误的结论,用户难以察觉。为此,本文提出基于预测性-可计算性-稳定性(PCS)框架的轻量级双重验证机制,通过合理扰动检测代理是否能可靠区分信号与噪声,作为结论可证伪性的约束,揭示结论是否建立在稳定信号之上、是否响应噪声或依赖输入偶然特征。我们在具有可控信噪比的合成数据上验证该方法,确认其能准确追踪真实信号强度。随后在11个真实数据集上应用该方法,使用OpenAI Codex分析,发现其中6个数据集的肯定结论缺乏充分支持,尽管单次运行可能呈现支持结果。进一步分析显示,ADS系统自报告置信度与其结论的实证稳定性严重不匹配。
原文摘要 · Abstract (English)
Agentic data science (ADS) pipelines have grown rapidly in both capability and adoption, with systems such as OpenAI Codex now able to directly analyze datasets and produce answers to statistical questions. However, these systems can reach falsely optimistic conclusions that are difficult for users to detect. To address this, we propose a pair of lightweight sanity checks grounded in the Predictability-Computability-Stability (PCS) framework for veridical data science. These checks use reasonable perturbations to screen whether an agent can reliably distinguish signal from noise, acting as a falsifiability constraint that can expose affirmative conclusions as unsupported. Together, the two checks characterize the trustworthiness of an ADS output, e.g. whether it has found stable signal, is responding to noise, or is sensitive to incidental aspects of the input. We validate the approach on synthetic data with controlled signal-to-noise ratios, confirming that the sanity checks track ground-truth signal strength. We then demonstrate the checks on 11 real-world datasets using OpenAI Codex, characterizing the trustworthiness of each conclusion and finding that in 6 of the datasets an affirmative conclusion is not well-supported, even though a single ADS run may support one. We further analyze failure modes of ADS systems and find that ADS self-reported confidence is poorly calibrated to the empirical stability of its conclusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。