用信息论衡量阅读时视觉输入质量,发现上半部分信息更重要。
Modeling Bottom-up Information Quality during Language Processing
- 用互信息衡量字形与词义的关联度,量化视觉输入质量。
- 中英文阅读实验显示,上半部分被遮挡时阅读更慢,且英语差异更明显。
- 适合研究语言认知、视觉信息处理或计算语言学的学者参考。
当代语言加工理论认为,理解过程同时依赖自上而下的预期和自下而上的输入。该模型预测:低质量的自下而上输入会导致理解困难。本文在阅读领域验证这一预测,提出以视觉信息与词义之间的互信息(MI)作为输入质量的信息论度量,并将此纳入贝叶斯阅读模型。通过在英语和汉语中比较遮挡单词上下半部分与完整单词的阅读时间,验证该度量。利用多模态语言模型估算视觉输入与词汇间的互信息,结果表明:遮挡上半部分比下半部分导致更显著的阅读延迟。此外,英语中上半部分的信息量显著高于下半部分,这种不对称性在阅读时间中得到体现。
原文摘要 · Abstract (English)
Contemporary theories model language processing as integrating both top-down expectations and bottom-up inputs. One major prediction of such models is that the quality of the bottom-up inputs modulates ease of processing -- noisy inputs should lead to difficult and effortful comprehension. We test this prediction in the domain of reading. First, we propose an information-theoretic operationalization for the "quality" of bottom-up information as the mutual information (MI) between visual information and word identity. We formalize this prediction in a mathematical model of reading as a Bayesian update. Second, we test our operationalization by comparing participants' reading times in conditions where words' information quality has been reduced, either by occluding their top or bottom half, with full words. We collect data in English and Chinese. We then use multimodal language models to estimate the mutual information between visual inputs and words. We use these data to estimate the specific effect of reduced information quality on reading times. Finally, we compare how information is distributed across visual forms. In English and Chinese, the upper half contains more information about word identity than the lower half. However, the asymmetry is more pronounced in English, a pattern which is reflected in the reading times.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。