arXiv:2602.16201cs.CLcs.AI2026-02被引 3

揭示大模型长尾知识缺失的根源与影响,提出系统性分析框架。

Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications

  • 构建四维分析框架:定义、机制、干预、影响,系统梳理长尾知识问题。
  • 指出现有评估方法掩盖了罕见但关键的知识失败,导致责任难追。
  • 适合关注模型公平性、可信度和治理的研究者与开发者参考。

大语言模型在网页规模语料上训练,知识分布呈现陡峭幂律,多数知识出现频率极低。尽管模型规模扩大提升了平均性能,但对低频、领域特定、文化及时间相关知识的持续失效仍缺乏清晰刻画。本文构建了长尾知识的结构化分类与分析体系,综合技术与社会技术视角的既有研究。提出一个四维分析框架:长尾知识的定义方式、训练与推理中知识丢失或扭曲的机制、缓解此类问题的技术干预措施,以及这些失效对公平性、可问责性、透明度与用户信任的影响。进一步分析现有评估实践如何掩盖尾部行为,使罕见但关键的失败难以追踪。论文最后指出了隐私、可持续性与治理等开放挑战,限制了长尾知识的有效表达。整体上,该工作为理解长尾知识的定义、流失、评估与部署表现提供了统一的概念框架。

原文摘要 · Abstract (English)

Large language models (LLMs) are trained on web-scale corpora that exhibit steep power-law distributions, in which the distribution of knowledge is highly long-tailed, with most appearing infrequently. While scaling has improved average-case performance, persistent failures on low-frequency, domain-specific, cultural, and temporal knowledge remain poorly characterized. This paper develops a structured taxonomy and analysis of long-Tail Knowledge in large language models, synthesizing prior work across technical and sociotechnical perspectives. We introduce a structured analytical framework that synthesizes prior work across four complementary axes: how long-Tail Knowledge is defined, the mechanisms by which it is lost or distorted during training and inference, the technical interventions proposed to mitigate these failures, and the implications of these failures for fairness, accountability, transparency, and user trust. We further examine how existing evaluation practices obscure tail behavior and complicate accountability for rare but consequential failures. The paper concludes by identifying open challenges related to privacy, sustainability, and governance that constrain long-Tail Knowledge representation. Taken together, this paper provides a unifying conceptual framework for understanding how long-Tail Knowledge is defined, lost, evaluated, and manifested in deployed language model systems.

长尾知识模型公平性评估漏洞可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。