对比大模型与人类对故事悬念感知的差异
Do Language Models Agree with Human Perceptions of Suspense in Stories?
- 用大模型替代人类复现经典悬念心理实验
- 模型能识别悬念存在但无法判断强度变化
- 揭示模型与人类在悬念感知上的本质不同
悬念是人类对叙事文本产生的情感反应,被认为涉及复杂的认知过程。我们复现了四项经典的关于人类悬念感知的心理学研究,将人类反应替换为不同开源与闭源大语言模型的响应。结果表明,虽然大模型能够判断文本是否旨在引发悬念,但无法准确估计文本序列中悬念的相对强度,也无法正确捕捉悬念在多个文本段落间的起伏变化。通过对抗性打乱故事文本,我们探究了模型与人类悬念感知差异的成因。结论是:尽管大模型能表面识别并追踪悬念的某些特征,但其处理悬念的方式与人类读者存在本质区别。
原文摘要 · Abstract (English)
Suspense is an affective response to narrative text that is believed to involve complex cognitive processes in humans. Several psychological models have been developed to describe this phenomenon and the circumstances under which text might trigger it. We replicate four seminal psychological studies of human perceptions of suspense, substituting human responses with those of different open-weight and closed-source LMs. We conclude that while LMs can distinguish whether a text is intended to induce suspense in people, LMs cannot accurately estimate the relative amount of suspense within a text sequence as compared to human judgments, nor can LMs properly capture the human perception for the rise and fall of suspense across multiple text segments. We probe the abilities of LM suspense understanding by adversarially permuting the story text to identify what cause human and LM perceptions of suspense to diverge. We conclude that, while LMs can superficially identify and track certain facets of suspense, they do not process suspense in the same way as human readers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。