大模型在填空任务中表现与人类差异显著,不能替代人类认知。
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
- 用填空任务对比大模型与人类生成行为
- 模型低估人类高频回答,高估罕见答案,语义空间差异大
- 适合研究语言模型偏差与认知建模的学者
本研究通过填空任务比较多个语言模型在下一个词预测层面的生成行为与人类表现。结果发现,尽管训练时间更长的大型模型通常更接近人类产出,但仍持续低估人类响应的概率,对罕见回答排名过高,对高频回答排名过低,且生成的语义空间显著不同于人类。这表明在可解释的领域内,语言模型的生成无法作为填空任务的替代或人类行为的模型。
原文摘要 · Abstract (English)
In this work we compare the generative behavior at the next token prediction level in several language models by comparing them to human productions in the cloze task. We find that while large models trained for longer are typically better estimators of human productions, but they reliably under-estimate the probabilities of human responses, over-rank rare responses, under-rank top responses, and produce highly distinct semantic spaces. Altogether, this work demonstrates in a tractable, interpretable domain that LM generations can not be used as replacements of or models of the cloze task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。