arXiv:2511.03070cs.AIcs.LG2025-11被引 3

测试大模型对真实世界概率分布的理解能力,发现其表现不佳。

Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge

  • 构建首个基准测试大模型对真实世界统计分布的掌握程度。
  • 模型在经济、健康等领域的分布预测准确率普遍偏低。
  • 结果表明模型缺乏观测层面知识,制约了因果推理能力。

人工智能系统在多个科学领域展现出巨大潜力,并被广泛应用于现实场景。尽管取得显著进展,仍需提升以实现更通用的智能。关键区别在于:事实性知识可验证真伪(如“英国首都是哪里?”),而概率性知识反映现实世界的概率特性(如“美国计算机科学毕业生的性别分布如何?”)。本文旨在建立首个基准,评估大语言模型(LLM)对真实世界人口分布的概率知识掌握情况。由于训练数据量庞大,推测模型可能内化部分真实分布。然而,经典统计学中的高维诅咒表明,在高维空间学习分布存在根本性挑战。本研究首次直接检验该假设,评估模型在经济学、健康、教育和社会行为等领域对实证分布的掌握能力。结果显示,模型整体表现较差,未自然内化真实世界统计数据。结合佩尔的因果层次理论(PCH),该基准表明语言模型缺乏观测分布知识(PCH第1层),因此其干预性(第2层)和反事实性(第3层)知识也受限。

原文摘要 · Abstract (English)

Artificial intelligence (AI) systems hold great promise for advancing various scientific disciplines, and are increasingly used in real-world applications. Despite their remarkable progress, further capabilities are expected in order to achieve more general types of intelligence. A critical distinction in this context is between factual knowledge, which can be evaluated against true or false answers (e.g., "what is the capital of England?"), and probabilistic knowledge, reflecting probabilistic properties of the real world (e.g., "what is the sex of a computer science graduate in the US?"). In this paper, our goal is to build a benchmark for understanding the capabilities of LLMs in terms of knowledge of probability distributions describing the real world. Given that LLMs are trained on vast amounts of text, it may be plausible that they internalize aspects of these distributions. Indeed, LLMs are touted as powerful universal approximators of real-world distributions. At the same time, classical results in statistics, known as curse of dimensionality, highlight fundamental challenges in learning distributions in high dimensions, challenging the notion of universal distributional learning. In this work, we develop the first benchmark to directly test this hypothesis, evaluating whether LLMs have access to empirical distributions describing real-world populations across domains such as economics, health, education, and social behavior. Our results demonstrate that LLMs perform poorly overall, and do not seem to internalize real-world statistics naturally. When interpreted in the context of Pearl's Causal Hierarchy (PCH), our benchmark demonstrates that language models do not contain knowledge on observational distributions (Layer 1 of PCH), and thus the Causal Hierarchy Theorem implies that interventional (Layer 2) and counterfactual (Layer 3) knowledge of these models is also limited.

大模型概率分布因果推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。