用合成数据和对抗验证,让自动科研代码的因果错误显性化
Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines

- 将研究问题转为结构化因果协议,生成带真实效应的模拟数据
- 在假设被破坏时检测出分析失效,而非悄悄输出错误结果
- 适合医学等敏感数据场景,帮助发现隐藏的因果漏洞
自动化研究系统虽能加速实证分析,却易出现无声失败:代码运行成功但基于无效因果假设。我们提出人工智能流行病学研究助手(ARA),通过显式编码因果设计原则、研究特定假设与方法约束,使这类错误可见。ARA将协议构建、合成数据生成与对抗验证整合为统一流程。它先将自然语言研究问题转化为结构化因果协议,再利用已知真实效应的结构因果模型(SCM)生成合成数据,支持在无法获取保密数据(如医疗数据)时开发分析流程。随后在受控地违反识别假设条件下评估生成分析。在自动化因果推理基准上评估其对识别策略、因果量、处理变量与结果变量的恢复能力,以及代码与协议的一致性。尽管协议构建与对抗验证未显著提升数值结果与基准估计的一致性,但改变了失败模式:不再静默返回因果估计,而是暴露协议问题、诊断失败、推断不完整或降级非因果解释。结果表明,有效性优先的自动化科学系统应不仅以答案准确率评估,更应考察其能否指出因果主张不成立的情况。
原文摘要 · Abstract (English)
While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes successfully yet relies on invalid causal assumptions. We present the Artificial Intelligence (AI)-based Epidemiology Research Assistant (ARA), a framework that makes these failures visible by explicitly encoding causal design principles, study-specific assumptions, and methodological constraints. ARA integrates protocol construction, synthetic data generation, and adversarial validation into a unified pipeline. The framework translates natural language research questions into structured causal protocols and executable analysis code by first constructing a protocol and then generating synthetic datasets using Structural Causal Models (SCMs) with known ground-truth effects. This synthetic-data step can also support pipeline development when access to confidential data, such as medical data, is restricted. The generated analysis is then evaluated under controlled violations of identification assumptions. We evaluate ARA on the Automated Causal Reasoning Benchmark, assessing recovery of identification strategies, causal quantities, treatment and outcome variables, and consistency between generated code and approved protocol. Protocol construction and adversarial validation did not consistently improve numerical agreement with benchmark estimates compared with standard LLM-based generation. However, they changed the failure mode: instead of silently returning causal estimates, ARA often surfaced protocol concerns, diagnostic failures, incomplete inference, or downgraded non-causal interpretations. These findings suggest that validity-first automated science systems should be evaluated not only by answer accuracy, but also by whether they indicate when causal claims are unwarranted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。