arXiv:2510.07231cs.CLcs.AI2025-10

测试大模型在不同经济情境下推理因果关系的能力,发现模型易受误导且难以调整判断。

EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models

  • 构建10,490个带情境标注的因果三元组数据集,源自顶级期刊实证研究。
  • 模型在需跨情境反转判断时准确率从73.9%降至41.3%,误导信息下低于50%。
  • 模型过度倾向给出方向性结论,对零效应识别率仅13.8%,校准差。

社会经济因果效应高度依赖制度与环境背景,同一干预在不同监管体系、市场条件、时间或人群下可能产生相反效果。这对大语言模型在决策支持中的表现提出挑战:能否根据指定情境推断因果方向,并在情境变化时修正判断?为此,我们构建EconCausal,一个包含10,490个带情境标注的因果三元组的大规模基准,源自2,595篇顶级经济与金融期刊的高质量实证研究,通过四阶段严格流程(多轮共识、上下文精炼、多评阅过滤)生成。实验显示,尽管顶尖模型在固定明确情境下可达88%准确率,但在需跨情境反转判断的任务中,准确率下降32.6~个百分点(73.9%→41.3%),引入误导性符号证据后进一步跌破50%。模型还过度坚持方向性判断,对零效应识别率仅为13.8%,在该类别上校准严重不足。数据集与基准已公开于https://anonymous.4open.science/r/econcausal-benchmark-6F12。

原文摘要 · Abstract (English)

Socio-economic causal effects depend heavily on their institutional and environmental contexts. The same intervention can produce different, even opposite, effects across regulatory regimes, market conditions, time periods, or populations. This poses a challenge for large language models (LLMs) in decision-support roles: can they infer the direction of a causal effect under a specified context, and revise that judgment when the context changes? To address this, we introduce EconCausal, a large-scale benchmark of 10,490 context-annotated causal triplets extracted from 2,595 high-quality empirical studies in top-tier economics and finance journals, constructed through a rigorous four-stage pipeline with multi-run consensus, context refinement, and multi-critic filtering. Across models, LLMs often fail to condition their predictions on context. While top models reach 88% accuracy in fixed, explicit contexts, accuracy falls by 32.6~pp on cases that require revising the sign across contexts (73.9% to 41.3%), and drops below 50% once misleading signed evidence is introduced. Models also over-commit to directional (+/-) signs, recognizing null effects only 13.8% of the time while remaining poorly calibrated on these categories. The dataset and benchmark are publicly available at https://anonymous.4open.science/r/econcausal-benchmark-6F12.

因果推理大模型评估经济分析情境感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。