从数据中自动发现治疗的未知因果效应,突破传统实验局限。
Exploratory Causal Inference in SAEnce
- 用预训练模型提取试验数据特征,再通过稀疏自编码器解释。
- 在生态学真实实验中首次实现无监督因果效应识别。
- 提出递归分层方法解决多重检验与效应混淆问题,适合科研人员
随机对照试验是科学的基石,但依赖人工假设和昂贵分析,难以大规模开展,可能局限于流行却不完整的假设。本文提出直接从数据中发现治疗的未知效应。通过预训练基础模型将试验的非结构化数据转化为有意义表征,并利用稀疏自编码器进行解释。然而,在神经层面发现显著因果效应面临多重检验问题和效应纠缠挑战。为此,我们引入神经效应搜索(Neural Effect Search),一种新颖的递归过程,通过渐进式分层解决上述问题。在半合成实验中验证算法稳健性后,我们在实验生态学背景下,首次成功实现了真实科学试验中的无监督因果效应识别。
原文摘要 · Abstract (English)
Randomized Controlled Trials are one of the pillars of science; nevertheless, they rely on hand-crafted hypotheses and expensive analysis. Such constraints prevent causal effect estimation at scale, potentially anchoring on popular yet incomplete hypotheses. We propose to discover the unknown effects of a treatment directly from data. For this, we turn unstructured data from a trial into meaningful representations via pretrained foundation models and interpret them via a sparse autoencoder. However, discovering significant causal effects at the neural level is not trivial due to multiple-testing issues and effects entanglement. To address these challenges, we introduce Neural Effect Search, a novel recursive procedure solving both issues by progressive stratification. After assessing the robustness of our algorithm on semi-synthetic experiments, we showcase, in the context of experimental ecology, the first successful unsupervised causal effect identification on a real-world scientific trial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。