用人类语言生成与理解的差异测试大模型认知合理性
Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility
- 以人类对因果动词中代词生成与理解的不对称性为基准
- 大模型在较大规模下更接近人类模式,且依赖提示设计
- 适合研究大模型语言认知机制的学者参考
大语言模型(LLMs)是否像人类一样处理语言,是理论与实践中的核心争议。本文从人类句子处理中的生成-理解差异出发,评估指令微调后的LLM是否复现这一现象。以人类在隐含因果动词情境下代词生成与理解的实证不对称性为测试基准,发现部分LLM在数量和质量上均表现出类人不对称性。该行为取决于模型规模——越大越接近人类模式,以及元语言提示的选择。实验代码与结果已公开于https://github.com/LingMechLab/Production-Interpretation_Asymmetries_ACL2025。
原文摘要 · Abstract (English)
Whether large language models (LLMs) process language similarly to humans has been the subject of much theoretical and practical debate. We examine this question through the lens of the production-interpretation distinction found in human sentence processing and evaluate the extent to which instruction-tuned LLMs replicate this distinction. Using an empirically documented asymmetry between pronoun production and interpretation in humans for implicit causality verbs as a testbed, we find that some LLMs do quantitatively and qualitatively reflect human-like asymmetries between production and interpretation. We demonstrate that whether this behavior holds depends upon both model size-with larger models more likely to reflect human-like patterns and the choice of meta-linguistic prompts used to elicit the behavior. Our codes and results are available at https://github.com/LingMechLab/Production-Interpretation_Asymmetries_ACL2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。