arXiv:2510.18368cs.CL2025-10被引 1

评测韩语大模型事实性,发现最强模型仅33.7%正确。

KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs

  • 构建1000个韩语事实问答题,答案明确易评分。
  • 最强模型准确率仅33.7%,远低于英语基准表现。
  • 揭示推理能力有助于模型调用知识并识别不确定项。

我们提出韩国简单问答基准(Korean SimpleQA, KoSimpleQA),用于评估大语言模型在韩语文化知识方面的真实性。该基准包含1000个简短、指向明确的事实类问题,答案无歧义,易于评分。我们在多个支持韩语的开源大模型上进行了全面评估,发现即使最强模型也仅能正确回答33.7%的问题,凸显了该数据集的挑战性。值得注意的是,模型在KoSimpleQA上的排名与英语SimpleQA显著不同,表明其独特价值。此外,对推理型模型的分析显示,引入推理机制可帮助模型更有效地激活潜在知识,并在不确定时选择不回答,提升可靠性。数据集已公开于 https://anonymous.4open.science/r/KoSimpleQA-62EB。

原文摘要 · Abstract (English)

We present $\textbf{Korean SimpleQA (KoSimpleQA)}$, a benchmark for evaluating factuality in large language models (LLMs) with a focus on Korean cultural knowledge. KoSimpleQA is designed to be challenging yet easy to grade, consisting of 1,000 short, fact-seeking questions with unambiguous answers. We conduct a comprehensive evaluation across a diverse set of open-source LLMs of varying sizes that support Korean, and find that even the strongest model generates correct answer only 33.7% of the time, underscoring the challenging nature of KoSimpleQA. Notably, performance rankings on KoSimpleQA differ substantially from those on the English SimpleQA, highlighting the unique value of our dataset. Furthermore, our analysis of reasoning LLMs shows that engaging reasoning capabilities in the factual QA task can both help models better elicit their latent knowledge and improve their ability to abstain when uncertain. KoSimpleQA can be found at https://anonymous.4open.science/r/KoSimpleQA-62EB.

事实性评测韩语LLM推理能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。