arXiv:2609.04835cs.CL2026-09

评估大模型不能只看对错,还要看它能否提供多样化的知识解答。

On Epistemic Diversity in Large Language Models

  • 提出用认知多样性衡量大模型的知识呈现范围
  • 发现顶尖模型常将多种有效答案压缩成少数标准答案
  • 适合关注模型思维广度与教育应用的研究者

大语言模型不仅用于信息检索,还承担回答问题、解释和教学等任务。在此类场景中,仅靠准确率或对齐度无法全面评估模型。一个模型可能给出正确答案,却限制用户接触其他有效答案、解释或推理路径。基于哲学与社会认识论中的认知多样性概念,本文将其形式化为大模型向用户展示的有效答案、解释和推理路径的多样性。我们提出初步框架以概念化和测量大模型的认知多样性,并在两个领域进行实证操作。结果发现,前沿大模型常表现出认知狭窄,反复将庞大的有效答案空间压缩至少数标准子集。这表明大模型评估应超越单纯准确性,将认知多样性视为重要能力维度。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A system may give a correct answer while still narrowing users' access %to knowledge. to alternative valid answers, explanations, or reasoning routes. Drawing on the broader notion of epistemic diversity in philosophy and social epistemology, we formalize it in the context of LLMs as the range of valid answers, explanations, and reasoning routes that an LLM exposes to users. We argue that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks. We propose a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs, and operationalize it in two domains. We find that frontier LLMs often exhibit epistemic narrowness, repeatedly collapsing large valid answer spaces onto small canonical subsets. These findings suggest that LLM evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as an important dimension of model capability.

大模型评估认知多样性知识呈现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。