用大模型知识指导因果发现中的干预实验设计,提升效率与准确性。
Can Large Language Models Help Experimental Design for Causal Discovery?
- 利用大模型的世界知识辅助选择最优干预目标
- 在4个真实基准上优于现有方法,甚至超越人类表现
- 适合需要高效实验设计的科研人员和自动化发现系统
在科学或因果发现中,设计合适的实验并选择最优干预目标是一个长期难题。仅从观测数据中识别潜在因果结构本身就很困难,而获取干预数据对因果发现至关重要,但通常成本高、耗时长。以往方法常依赖不确定性或梯度信号来确定干预目标,但在初始阶段干预数据有限时,这些数值信号估计不准确,可能导致次优结果。本文探索一种新方法:利用大语言模型(LLMs)中蕴含的丰富实验设计知识,辅助因果发现中的干预目标选择。我们提出大型语言模型引导的干预目标选择框架(LeGIT),该框架能有效融合大模型知识,增强现有数值方法在干预目标选择上的性能。在4个真实基准上,LeGIT表现出显著改进和更强鲁棒性,甚至超过人类表现,证明了大模型在辅助科学发现实验设计中的有效性。
原文摘要 · Abstract (English)
Designing proper experiments and selecting optimal intervention targets is a longstanding problem in scientific or causal discovery. Identifying the underlying causal structure from observational data alone is inherently difficult. Obtaining interventional data, on the other hand, is crucial to causal discovery, yet it is usually expensive and time-consuming to gather sufficient interventional data to facilitate causal discovery. Previous approaches commonly utilize uncertainty or gradient signals to determine the intervention targets. However, numerical-based approaches may yield suboptimal results due to the inaccurate estimation of the guiding signals at the beginning when with limited interventional data. In this work, we investigate a different approach, whether we can leverage Large Language Models (LLMs) to assist with the intervention targeting in causal discovery by making use of the rich world knowledge about the experimental design in LLMs. Specifically, we present Large Language Model Guided Intervention Targeting (LeGIT) -- a robust framework that effectively incorporates LLMs to augment existing numerical approaches for the intervention targeting in causal discovery. Across 4 realistic benchmark scales, LeGIT demonstrates significant improvements and robustness over existing methods and even surpasses humans, which demonstrates the usefulness of LLMs in assisting with experimental design for scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。