arXiv:2510.04226cs.CLcs.AI2025-10中稿 · EMNLP被引 15

首次系统评估大模型知识多样性,发现进步但仍有明显不足。

What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models

  • 跨时间跨文化测试27个大模型,覆盖155个主题和12国数据。
  • 过去三年知识多样性显著提升,但所有模型仍不如搜索引擎多样。
  • 小模型反而比大模型更丰富,英语知识远超本地语言知识。

大型语言模型(LLMs)日益成为主要知识来源,但其表征的现实主张多样性——即知识多样性——从未被量化测量。低知识多样性可能导致知识萎缩风险,使同质化模型随时间限制信息获取范围。当前主流观点认为整体多样性偏低,但均基于单一时间点,缺乏基准参考且忽略国家间差异。本文首次系统研究了跨时间和文化语境下的知识多样性,对27个模型在155个主题上进行测试,覆盖12个国家,共生成170万条回应和7000万条独立主张。结果表明,过去三年知识多样性显著上升,扭转了近期的悲观预期。然而,所有系统仍低于搜索基线。该差距并非均匀:检索增强生成(RAG)可提升多样性,而大模型反而比小模型更缺乏多样性。此外,模型参数知识系统性地偏向英语,远超本地语言知识。综合来看,尽管知识多样性有所进步,但仍不足且分布不均。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as primary knowledge sources, yet their epistemic diversity - defined as the diversity of real-world claims in their outputs - has never been measured. Low epistemic diversity would pose a risk of knowledge collapse as homogeneous LLMs mediate a shrinking in the range of accessible information over time. The dominant paradigm is that overall LLM diversity is low, but this is always with respect to a single point in time, with no reference baseline or consideration for variation across countries. We address this gap in knowledge by performing the first systematic study of epistemic diversity in LLMs across time and cultural context, testing 27 LLMs on 155 topics covering 12 countries, resulting in 1.7M responses and 70M individual claims. We find that epistemic diversity has increased substantially over the past three years, a positive counter to recent diversity pessimism. However, despite progress, we find that every system is less diverse than a search baseline. This gap is not uniform: RAG can improve diversity, while large models are counterintuitively less diverse than smaller ones. Moreover, LLM parametric knowledge systematically reflects English over local-language knowledge for country specific topics. Together, these results demonstrate that while progress on epistemic diversity is tangible, it is insufficient and unevenly distributed.

知识多样性大模型评估跨文化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。