大模型更准预测人类填空答案,因语义理解更强但不依赖词汇搭配。
On the scaling relationship between cloze probabilities and language model next-token prediction
- 用填空任务测试模型预测下一个词的能力,比较不同规模模型表现。
- 大模型对人类回答的预测概率更高,且更符合语义合理性。
- 适合研究语言模型认知机制或人类语言行为建模的研究者。
近期研究表明,更大的语言模型在预测眼动和阅读时间数据方面具有更强的预测能力。尽管最佳模型对人类回答的概率分配仍不足,但更大模型在填空数据中对下一词及其生成可能性的估计质量更高,因其对词汇共现统计的敏感度较低,而与人类填空回答的语义对齐更好。结果支持了大模型更强的记忆容量有助于其猜测更语义恰当的词语,但降低了对单词识别相关低层次信息的敏感性。
原文摘要 · Abstract (English)
Recent work has shown that larger language models have better predictive power for eye movement and reading time data. While even the best models under-allocate probability mass to human responses, larger models assign higher-quality estimates of next tokens and their likelihood of production in cloze data because they are less sensitive to lexical co-occurrence statistics while being better aligned semantically to human cloze responses. The results provide support for the claim that the greater memorization capacity of larger models helps them guess more semantically appropriate words, but makes them less sensitive to low-level information that is relevant for word recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。