arXiv:2511.12300cs.CL2025-11

对比大模型与人类在日语问答中的难度差异,发现两者难易点不一致。

Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering

  • 用真实日语抢答题数据,对比大模型与人类答题正确率
  • 大模型在无维基覆盖答案和需数值回答的题目上表现更差
  • 适合研究模型认知偏差或人机差异的从业者参考

大语言模型在诸多自然语言处理任务中已超越人类表现,但尚不清楚对人类而言困难的问题是否同样难倒大模型。本研究探究了在抢答场景下,日语问答题对大模型与人类而言的难度差异。首先,收集包含问题、答案及人类正确率的日语问答数据;随后,在多种提示设置下让大模型作答,并从两个分析视角比较其正确率与人类表现。实验结果表明,相较于人类,大模型在答案未被维基百科覆盖的题目上表现更差,且在需要数值回答的问题上也存在显著困难。

原文摘要 · Abstract (English)

LLMs have achieved performance that surpasses humans in many NLP tasks. However, it remains unclear whether problems that are difficult for humans are also difficult for LLMs. This study investigates how the difficulty of quizzes in a buzzer setting differs between LLMs and humans. Specifically, we first collect Japanese quiz data including questions, answers, and correct response rate of humans, then prompted LLMs to answer the quizzes under several settings, and compare their correct answer rate to that of humans from two analytical perspectives. The experimental results showed that, compared to humans, LLMs struggle more with quizzes whose correct answers are not covered by Wikipedia entries, and also have difficulty with questions that require numerical answers.

大模型评估人机对比日语NLP问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。