验证了语言模型能否稳定复现人类对句子的可预测性度量。
Can LLMs capture stable human-generated sentence entropy measures?
- 通过自助法分析发现,90%句子在111次(德语)或81次(英语)响应内收敛
- 高可预测性句子仅需20次响应,而高熵句子需更多样本才能稳定
- GPT-4o最接近人类数据,但结果依赖提取方法和提示设计
预测下一个词是语言理解的核心机制,可用香农熵量化。目前尚无实证共识说明需要多少人类响应才能获得稳定的词级熵估计。大型语言模型(LLMs)正被用作人类规范数据的替代品,但其复现稳定人类熵的能力仍不明确。本文利用两个公开的德语和英语填空数据集,采用基于自助法的收敛分析,追踪熵估计随样本量变化的稳定性。两种语言中,超过97%的句子在可用样本量内达到稳定。德语90%句子在111次响应后收敛,英语为81次;低熵句子(<1)仅需20次,高熵句子(>2.5)则需更多。研究首次直接验证了常见规范做法,并表明收敛程度取决于句子可预测性。随后将稳定的人类熵与多个LLM(包括GPT-4o、GPT2-xl/german-GPT-2、RoBERTa Base/GottBERT、LLaMA 2 7B Chat)的熵估计对比,使用基于逻辑值的概率提取和基于采样的频率估计。GPT-4o表现最佳,但一致性高度依赖提取方法与提示设计:逻辑值估计误差更小,采样估计更能捕捉人类变异分布。结果为人类规范提供实用指南,表明尽管LLM可近似人类熵,但不能取代稳定的人类分布。
原文摘要 · Abstract (English)
Predicting upcoming words is a core mechanism of language comprehension and may be quantified using Shannon entropy. There is currently no empirical consensus on how many human responses are required to obtain stable and unbiased entropy estimates at the word level. Moreover, large language models (LLMs) are increasingly used as substitutes for human norming data, yet their ability to reproduce stable human entropy remains unclear. Here, we address both issues using two large publicly available cloze datasets in German 1 and English 2. We implemented a bootstrap-based convergence analysis that tracks how entropy estimates stabilize as a function of sample size. Across both languages, more than 97% of sentences reached stable entropy estimates within the available sample sizes. 90% of sentences converged after 111 responses in German and 81 responses in English, while low-entropy sentences (<1) required as few as 20 responses and high-entropy sentences (>2.5) substantially more. These findings provide the first direct empirical validation for common norming practices and demonstrate that convergence critically depends on sentence predictability. We then compared stable human entropy values with entropy estimates derived from several LLMs, including GPT-4o, using both logit-based probability extraction and sampling-based frequency estimation, GPT2-xl/german-GPT-2, RoBERTa Base/GottBERT, and LLaMA 2 7B Chat. GPT-4o showed the highest correspondence with human data, although alignment depended strongly on the extraction method and prompt design. Logit-based estimates minimized absolute error, whereas sampling-based estimates were better in capturing the dispersion of human variability. Together, our results establish practical guidelines for human norming and show that while LLMs can approximate human entropy, they are not interchangeable with stable human-derived distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。