arXiv:2607.21445cs.CL2026-07

测试大模型在日常文化知识上的表现,发现其在流行文化上明显弱于历史地理等硬知识。

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

论文配图:When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs
图 1 · 摘自论文原文
  • 构建多语言问答基准TriviaRoomQA,覆盖288个日常话题
  • 模型在流行文化题上准确率显著低于历史地理类题目
  • 不同语言表现差异大,知识获取非完全语言独立

测验室、趣味问答夜和知识竞赛挑战人类在广泛主题上的知识,包括经典事实与日常文化。本文通过问答题评估大语言模型在这些场景中的表现,涵盖常见与冷门话题。我们引入TriviaRoomQA,一个面向288个主题的多语言基准,包含6种欧洲语言的3,300道平行多选题,以及额外5,340道法语专属题目用于细粒度分析。我们评估了来自欧、亚、北美厂商的30个开源权重大模型,参数规模从7到70B不等。结果表明,模型在历史、地理、数学等知识密集型任务上表现良好,但在名人、音乐、电影、新闻等日常流行文化题上显著薄弱。此外,相同问题在不同语言下的表现存在差异,表明事实知识获取并非总是语言无关。整体显示现有学术饱和基准未能捕捉这一重要知识缺口。

原文摘要 · Abstract (English)

Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, we examine whether large language models (LLMs) can perform competitively in such settings, using quiz-style questions to test them on both common and niche topics. We introduce TriviaRoomQA, a multilingual benchmark designed to evaluate everyday, culturally grounded, and long-tail knowledge across 288 topics. The benchmark contains 3,300 parallel multiple-choice questions in six European languages and additional 5,340 French-only questions for a more fine-grained case study. We evaluate 30 open-weight LLMs from European, Asian, and North American providers, covering models from 7 to 70B parameters. We find that models are strong on knowledge-intensive topics such as history, geography, and mathematics, but substantially weaker on everyday popular-culture topics such as celebrities, music, movies, and news. Moreover, model performance varies across languages even for the same underlying questions, suggesting that access to factual knowledge is not always language-independent. In sum, our dataset and experiments demonstrate an important knowledge gap which is not captured by existing academic-based saturated benchmarks.

大模型评测多语言日常知识流行文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。