arXiv:2506.05000cs.CL2025-06ACL被引 1

从认知视角评测大模型理解过程,发现其仍不靠谱。

SCOP: Evaluating the Comprehension Process of Large Language Models from a Cognitive View

  • 构建五项认知能力评估体系,系统检验模型理解过程
  • 实测显示模型在局部信息理解上优于全局,但易通过错误路径得正确答案
  • 适合关注模型可解释性与可信推理的研究者

尽管大语言模型在机器理解方面潜力巨大,但在真实场景中仍难以完全信赖。这可能是因为缺乏对模型理解过程是否与专家一致的合理解释。本文提出SCOP,从认知视角细致考察大模型的理解过程。该方法包含五个必要理解技能的系统定义、严格的数据构建框架,以及对先进开源与闭源模型的详细分析。结果表明,大模型仍难以达到专家级理解水平。尽管如此,模型在局部信息理解上表现优于全局信息。进一步分析发现,模型可能存在不可靠性——可能通过有缺陷的理解路径得出正确答案。基于此,建议未来改进应更关注理解过程本身,确保训练中全面培养各项理解能力。

原文摘要 · Abstract (English)

Despite the great potential of large language models(LLMs) in machine comprehension, it is still disturbing to fully count on them in real-world scenarios. This is probably because there is no rational explanation for whether the comprehension process of LLMs is aligned with that of experts. In this paper, we propose SCOP to carefully examine how LLMs perform during the comprehension process from a cognitive view. Specifically, it is equipped with a systematical definition of five requisite skills during the comprehension process, a strict framework to construct testing data for these skills, and a detailed analysis of advanced open-sourced and closed-sourced LLMs using the testing data. With SCOP, we find that it is still challenging for LLMs to perform an expert-level comprehension process. Even so, we notice that LLMs share some similarities with experts, e.g., performing better at comprehending local information than global information. Further analysis reveals that LLMs can be somewhat unreliable -- they might reach correct answers through flawed comprehension processes. Based on SCOP, we suggest that one direction for improving LLMs is to focus more on the comprehension process, ensuring all comprehension skills are thoroughly developed during training.

大模型理解认知评测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。