arXiv:2605.29631cs.CLcs.AI2026-05

用自然语言查询预测因果效应,提升医学与社科研究效率

Predicting Causal Effects from Natural Language Queries using Structured Representations

论文配图:Predicting Causal Effects from Natural Language Queries using Structured Representations
图 1 · 摘自论文原文
  • 先生成查询的结构化表示,再预测效应大小
  • 微调后误差降低27%至71%,显著优于直接提示大模型
  • 适合需要快速评估实验效果的研究者使用

随机对照试验是医学和社会科学中可靠估计因果效应的基石,但其成本高、耗时长,促使人们探索从已有实验证据中预测因果效应。大型语言模型在知识密集型任务中表现优异,引发了对其能否预测因果效应大小的思考。为此,我们提出了Query2Effect,一个包含超过72,000个自然语言问题与实验描述对的大规模基准,通过在隐含性、抽象性和模糊性维度上变化查询,模拟真实的检索场景。我们进一步提出一种两步框架:首先生成查询的合成结构化表示,然后使用监督编码器模型预测效应大小。实验表明,微调对性能提升至关重要,绝对误差相比未微调的提示式大模型降低27%至71%;且该两步框架有利于跨领域泛化,凸显了语义理解与数值估计分离的优势。

原文摘要 · Abstract (English)

Randomized controlled trials are a cornerstone of medicine and the social sciences as they enable reliable estimates of causal effects. However, they are costly and time-consuming to conduct, motivating interest in predicting causal effects from existing experimental evidence. Recent advances in large language models (LLMs) have demonstrated strong performance on knowledge-intensive tasks, raising the question of whether these models can be used for forecasting causal effect sizes. To investigate this, we introduce Query2Effect, a new large-scale benchmark consisting of more than 72,000 natural language questions aligned with experiment descriptions, created to simulate realistic information-seeking scenarios by varying query specificity along dimensions of implicitness, abstraction, and ambiguity. We then propose a two-step framework that first generates a synthetic structured representation of a query before predicting effect size using a supervised encoder model. Experiments show that finetuning plays a crucial role in improving prediction performance, with absolute error reducing by -27% up to -71% compared to prompted out-of-the-box LLMs, and that our two-step framework is beneficial for out-of-domain generalization, highlighting the benefits of separating semantic interpretation from numerical effect estimation.

因果推断大模型应用自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。