限制认知资源让大模型更像人理解句子。
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models

- 设计双任务实验,同时算数和理解句子,模拟人类资源限制。
- 模型在双任务下对合理句的识别准确率显著高于不合理句。
- 适合研究大模型认知机制或类人推理的学者参考。
语言模型(LMs)在认知资源受限时行为更接近人类,尤其在预测句子处理成本(如阅读时间)方面。然而,这种约束是否同样影响句子理解策略尚不明确。现有方法未能直接平衡记忆存储与句子处理之间的关系,而这是人类工作记忆的核心。为此,我们提出一种双任务范式,将算术计算任务与句子理解任务结合,例如“The 2 cocktail + blended 3 =...”。实验表明,在双任务条件下,GPT-4o、o3-mini 和 o4-mini 模型转向基于合理性的理解方式,与人类的理性推理一致。具体而言,这些模型在双任务中对合理句子(如“鸡尾酒被调酒师调制”)与不合理句子(如“调酒师被鸡尾酒调制”)的准确率差距显著大于单任务条件。结果表明,记忆与处理资源的平衡限制能促进语言模型产生理性推理。更广泛地,这支持了人类类似句意理解本质上源于有限认知资源分配的观点。
原文摘要 · Abstract (English)
Language models (LMs) behave more like humans when their cognitive resources are restricted, particularly in predicting sentence processing costs such as reading times. However, it remains unclear whether such constraints similarly affect sentence comprehension strategies. Besides, existing methods do not directly target the balance between memory storage and sentence processing, which is central to human working memory. To address this issue, we propose a dual-task paradigm that combines an arithmetic computation task with a sentence comprehension task, such as "The 2 cocktail + blended 3 =..." Our experiments show that under dual-task conditions, GPT-4o, o3-mini, and o4-mini shift toward plausibility-based comprehension, mirroring humans' rational inference. Specifically, these models show a greater accuracy gap between plausible sentences (e.g., "The cocktail was blended by the bartender") and implausible sentences (e.g., "The bartender was blended by the cocktail") in the dual-task condition compared to the single-task conditions. These findings suggest that constraints on the balance between memory and processing resources promote rational inference in LMs. More broadly, they support the view that human-like sentence comprehension fundamentally arises from the allocation of limited cognitive resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。