arXiv:2604.06201cs.CLcs.AI2026-04

测试大模型理解文本分布信息的能力,发现模型在趋势判断上表现参差。

Beyond Facts: Benchmarking Distributional Reading Comprehension in Large Language Models

论文配图:Beyond Facts: Benchmarking Distributional Reading Comprehension in Large Language Models
图 1 · 摘自论文原文
  • 基于真实YouTube评论构建分布式阅读理解评测集
  • 模型能准确估算正负评比例,但对不同分布类型表现差异大
  • 适合研究大模型在真实世界趋势分析中的应用

当前大多数大语言模型阅读理解评测聚焦于可通过定位文本证据回答的事实性问题,但许多实际任务需要理解分布性信息,如跨文本集合表达的人群趋势和偏好。本文提出Text2DistBench,一个评估大模型从自然语言中推断分布知识能力的阅读理解基准。该基准基于真实世界关于电影和音乐实体的YouTube评论构建,提供实体元数据及关联评论,要求模型回答分布性问题,如估计正面与负面评论的比例,或识别观众讨论中最常见和第二常见的主题。为支持可靠且长期的评估,Text2DistBench的构建流程完全自动化,并持续更新以纳入新出现的实体。在多个大模型上的实验表明,尽管模型显著优于随机基线,但在不同分布类型和特征下的表现差异明显。这些发现揭示了当前大模型在分布阅读理解中的能力和局限性,也证明了Text2DistBench作为未来研究实用且可扩展的测试平台的价值。

原文摘要 · Abstract (English)

While most reading comprehension benchmarks for LLMs focus on factual information that can be answered by localizing specific textual evidence, many real-world tasks require understanding distributional information, such as population-level trends and preferences expressed across collections of text. We introduce Text2DistBench, a reading comprehension benchmark for evaluating LLMs' ability to infer distributional knowledge from natural language. Built from real-world YouTube comments about movie and music entities, the benchmark provides models with entity metadata and associated comments, and requires them to answer distributional questions, such as estimating the proportions of positive and negative comments, or identifying the most and second most frequent topics discussed among viewers. To support reliable and long-term evaluation, the construction pipeline of Text2DistBench is fully automated and continuously updated to incorporate newly emerging entities over time. Experiments across multiple LLMs show that while models substantially outperform random baselines, performance varies widely across different distribution types and characteristics. These findings highlight both the capabilities and limitations of current LLMs in distributional reading comprehension and demonstrate the value of Text2DistBench as a practical and scalable testbed for future research.

阅读理解分布推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。