首个评估大模型多语言事实能力的基准,测记忆也测自知之明。
Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering
- 用单一知识点、唯一答案设计问题,支持大模型自评
- 覆盖9种语言,分通用与语言特有双领域,评估更全面
- 发现主流模型在不同领域表现差异大,适合多语言研究者使用
我们提出KoLasSimpleQA,首个评估大语言模型(LLM)多语言事实能力的基准。受已有研究启发,构建的问题集具有单知识点覆盖、绝对客观性、唯一答案和时间稳定性等特点,支持高效使用‘大模型作为裁判’范式,测试模型的事实记忆与自我认知能力(‘知道不知道’)。该基准在两个关键维度拓展现有研究:(1)广度(多语言覆盖):包含9种语言,支持全球适用性评估;(2)深度(双域设计):涵盖通用领域(全球常识)和语言特有领域(如历史、文化、地区传统),实现对多语言能力的综合评估。我们评估了主流大语言模型,包括传统大模型与新兴的大推理模型。结果表明,两类模型在双领域间表现存在显著差异,体现在性能指标、排名、校准性和鲁棒性等方面。这凸显了在多语言场景下针对性评估与优化的必要性。我们希望KoLasSimpleQA能帮助研究社区更准确识别大模型在多语言环境中的能力边界,并为模型优化提供指导。项目将开源发布于https://github.com/opendatalab/KoLasSimpleQA。
原文摘要 · Abstract (English)
We introduce KoLasSimpleQA, the first benchmark evaluating the multilingual factual ability of Large Language Models (LLMs). Inspired by existing research, we created the question set with features such as single knowledge point coverage, absolute objectivity, unique answers, and temporal stability. These questions enable efficient evaluation using the LLM-as-judge paradigm, testing both the LLMs' factual memory and self-awareness ("know what they don't know"). KoLasSimpleQA expands existing research in two key dimensions: (1) Breadth (Multilingual Coverage): It includes 9 languages, supporting global applicability evaluation. (2) Depth (Dual Domain Design): It covers both the general domain (global facts) and the language-specific domain (such as history, culture, and regional traditions) for a comprehensive assessment of multilingual capabilities. We evaluated mainstream LLMs, including traditional LLM and emerging Large Reasoning Models. Results show significant performance differences between the two domains, particularly in performance metrics, ranking, calibration, and robustness. This highlights the need for targeted evaluation and optimization in multilingual contexts. We hope KoLasSimpleQA will help the research community better identify LLM capability boundaries in multilingual contexts and provide guidance for model optimization. We will release KoLasSimpleQA at https://github.com/opendatalab/KoLasSimpleQA .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。