arXiv:2505.08996cs.CL2025-05被引 3

新实验显示大模型理解谜题语句能力不输人类,甚至更优。

A suite of LMs comprehend puzzle statements as well as humans

  • 通过限制重读模拟真实阅读,发现人类准确率降至73%
  • Falcon-180B-Chat和GPT-4分别达76%和81%准确率
  • 模型在互惠动作类问题上与人类有相似理解难点

近期研究声称大语言模型在理解极简英语陈述时表现不如人类(Dentella等,2024)。本文重新审视该结论,认为人类表现被高估,而大模型能力被低估。我们使用相同刺激材料,开展预注册研究,比较人类在两种条件下反应:一种允许重读(复现原研究),另一种禁止重读(更自然的阅读测试)。当禁止重读时,人类准确率显著下降至73%,低于Falcon-180B-Chat的76%和GPT-4的81%。最新GPT-o1模型实现完美准确率。结果还表明,人类与模型在涉及潜在互惠行为(如亲吻)的问题上均表现不佳,反映共享的语用敏感性而非模型缺陷。通过对Llama-2-70B日志概率、模型开放式回答再编码及句子语法性评分的额外分析,揭示了对模型性能的系统性低估。发现GPT-4o的语法判断可随提示框架切换至普通或专家水平。这些发现强调需更严谨的实验设计与标注实践,挑战当前模型语言理解弱于人类的假设。

原文摘要 · Abstract (English)

Recent claims suggest that large language models (LMs) underperform humans in comprehending minimally complex English statements (Dentella et al., 2024). Here, we revisit those findings and argue that human performance was overestimated, while LLM abilities were underestimated. Using the same stimuli, we report a preregistered study comparing human responses in two conditions: one allowed rereading (replicating the original study), and one that restricted rereading (a more naturalistic comprehension test). Human accuracy dropped significantly when rereading was restricted (73%), falling below that of Falcon-180B-Chat (76%) and GPT-4 (81%). The newer GPT-o1 model achieves perfect accuracy. Results further show that both humans and models are disproportionately challenged by queries involving potentially reciprocal actions (e.g., kissing), suggesting shared pragmatic sensitivities rather than model-specific deficits. Additional analyses using Llama-2-70B log probabilities, a recoding of open-ended model responses, and grammaticality ratings of other sentences reveal systematic underestimation of model performance. We find that GPT-4o can align with either naive or expert grammaticality judgments, depending on prompt framing. These findings underscore the need for more careful experimental design and coding practices in LLM evaluation, and they challenge the assumption that current models are inherently weaker than humans at language comprehension.

语言理解大模型评估认知实验人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。