arXiv:2409.06518cs.CLcs.AI2024-09被引 1

用奥运奖牌榜测试大模型,发现它记得住数字却排不好名。

Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings

  • 用奥运奖牌数据检验大模型知识组织方式
  • 模型能准确回忆奖牌数但无法正确排序
  • 适合研究大模型认知缺陷的学者参考

大型语言模型(LLMs)在自然语言处理任务中取得了显著成就,但其内部知识结构仍不清晰。本研究通过历史奥运奖牌榜来考察这些结构,评估了大模型在两个任务上的表现:(1) 检索特定国家的奖牌数量;(2) 识别各国排名。尽管当前最先进的大模型在回忆奖牌数量方面表现优异,但在提供正确排名时存在明显困难,这凸显了其知识组织方式与人类推理之间的关键差异。该研究揭示了大模型内部知识整合的局限性,并为改进方向提供了启示。为促进后续研究,我们公开了代码、数据集及模型输出结果。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable success in natural language processing tasks, yet their internal knowledge structures remain poorly understood. This study examines these structures through the lens of historical Olympic medal tallies, evaluating LLMs on two tasks: (1) retrieving medal counts for specific teams and (2) identifying rankings of each team. While state-of-the-art LLMs excel in recalling medal counts, they struggle with providing rankings, highlighting a key difference between their knowledge organization and human reasoning. These findings shed light on the limitations of LLMs' internal knowledge integration and suggest directions for improvement. To facilitate further research, we release our code, dataset, and model outputs.

大模型评测知识结构奥运数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。