arXiv:2603.15542cs.CYcs.AI2026-03被引 2

评测大模型在真实社会政策干预中的因果推理能力,发现现有模型表现不佳。

InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems

  • 基于744项真实社会科学研究构建推理任务,无需预设因果图
  • 顶尖大模型在真实干预场景中推理准确率不足,暴露能力短板
  • 提出多智能体框架STRIDES,显著提升模型因果推理性能

社会科学研究中的因果推断依赖于以干预为中心、贯穿全程的研究设计推理,但当前基准无法评估大语言模型(LLMs)在该方面的能力。本文提出InterveneBench,一个面向真实社会系统中干预推理与因果研究设计的基准。每个任务均源自经过同行评审的社会科学实证研究,要求模型在未提供预定义因果图或结构方程的情况下,对政策干预和识别假设进行推理。InterveneBench涵盖744项跨政策领域的研究。实验结果表明,当前最先进大模型在此设定下表现不佳。为应对这一局限,我们进一步提出多智能体框架STRIDES,其在多项指标上显著优于现有推理模型。代码与数据已公开于https://github.com/Sii-yuning/STRIDES。

原文摘要 · Abstract (English)

Causal inference in social science relies on end-to-end, intervention-centered research-design reasoning grounded in real-world policy interventions, but current benchmarks fail to evaluate this capability of large language models (LLMs). We present InterveneBench, a benchmark designed to assess such reasoning in realistic social settings. Each instance in InterveneBench is derived from an empirical social science study and requires models to reason about policy interventions and identification assumptions without access to predefined causal graphs or structural equations. InterveneBench comprises 744 peer-reviewed studies across diverse policy domains. Experimental results show that state-of-the-art LLMs struggle under this setting. To address this limitation, we further propose a multi-agent framework, STRIDES. It achieves significant performance improvements over state-of-the-art reasoning models. Our code and data are available at https://github.com/Sii-yuning/STRIDES.

因果推理大模型评测政策分析多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。