arXiv:2412.18120cs.CLcs.AI2024-12被引 3

测试语言模型认知能力时,表现差可能因不懂任务而非记忆差。

Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm

  • 用不同难度的N-back任务检验模型对指令的理解能力
  • 高难度任务下性能下降,显示任务理解存在瓶颈
  • 适合研究大模型认知评估方法的学者参考

认知任务最初为人类设计,现被广泛用于评估语言模型。尽管应用看似直接,但结果解读常存困惑:模型表现不佳,究竟是认知能力不足,还是根本没理解任务?有研究认为GPT 3.5在2-back和3-back任务中表现下降,反映其工作记忆容量类似人类(Gong et al., 2024)。我们分析了多个开源语言模型在该任务上的表现,发现性能低下至少部分源于任务理解与任务状态维持的局限。通过引入更难的10-back任务、尝试不同提示策略,并分析模型注意力机制,我们进一步揭示了理解偏差的存在。本研究旨在推动语言模型认知评估方法的优化讨论。

原文摘要 · Abstract (English)

Cognitive tasks originally developed for humans are now increasingly used to study language models. While applying these tasks is often straightforward, interpreting their results can be challenging. In particular, when a model underperforms, it is often unclear whether this results from a limitation in the cognitive ability being tested or a failure to understand the task itself. A recent study argues that GPT 3.5's declining performance on 2-back and 3-back tasks reflects a working memory capacity limit similar to humans (Gong et al., 2024). By analyzing a range of open-source language models of varying performance levels on these tasks, we show that the poor performance is due at least in part to a limitation in task comprehension and task set maintenance. We challenge the best-performing model with progressively harder versions of the task (up to 10-back) and experiment with alternative prompting strategies, before analyzing model attentions. Our larger aim is to contribute to the ongoing conversation around refining methodologies for the cognitive evaluation of language models.

认知评估语言模型N-back

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。