arXiv:2508.10581cs.LG2025-08

用大模型助手帮非专业人士做因果推断,降低门槛。

Technical Report: Facilitating the Adoption of Causal Inference Methods Through LLM-Empowered Co-Pilot

  • 用大模型构建因果图并自动确定调整变量
  • 提出最小不确定性调整集标准,提升估计稳健性
  • 适合医疗、政策等领域研究者快速上手因果分析

从观察数据中估计处理效应(TE)在医疗、经济和公共政策等领域至关重要,但因需掌握复杂的因果假设、调整策略和模型选择而难以推广。本文提出 CATE-B,一个开源的智能协作者系统,利用大语言模型(LLMs)在代理框架下指导用户完成治疗效应估计的全流程:(i) 通过因果发现与大模型驱动的边定向构建结构因果模型;(ii) 基于新颖的最小不确定性调整集准则识别鲁棒的调整集;(iii) 根据因果结构和数据特征选择合适的回归方法。为促进可复现性和评估,我们发布了涵盖多种领域和因果复杂度的基准任务套件。通过融合因果推断与智能交互支持,CATE-B 降低了严谨因果分析的门槛,并为自动化处理效应估计建立了新范式。

原文摘要 · Abstract (English)

Estimating treatment effects (TE) from observational data is a critical yet complex task in many fields, from healthcare and economics to public policy. While recent advances in machine learning and causal inference have produced powerful estimation techniques, their adoption remains limited due to the need for deep expertise in causal assumptions, adjustment strategies, and model selection. In this paper, we introduce CATE-B, an open-source co-pilot system that uses large language models (LLMs) within an agentic framework to guide users through the end-to-end process of treatment effect estimation. CATE-B assists in (i) constructing a structural causal model via causal discovery and LLM-based edge orientation, (ii) identifying robust adjustment sets through a novel Minimal Uncertainty Adjustment Set criterion, and (iii) selecting appropriate regression methods tailored to the causal structure and dataset characteristics. To encourage reproducibility and evaluation, we release a suite of benchmark tasks spanning diverse domains and causal complexities. By combining causal inference with intelligent, interactive assistance, CATE-B lowers the barrier to rigorous causal analysis and lays the foundation for a new class of benchmarks in automated treatment effect estimation.

因果推断大模型应用智能辅助处理效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。