用隐含因果动词测试人类与大模型的语篇偏见,评估模型理解能力。
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
- 通过对比人类与多语言大模型在隐含因果动词上的反应,分析语篇偏见
- 仅最大规模德语模型(Bloom 6.4B)在指代关系上表现类人偏见
- 发现大模型在解释性关联和指称表达上缺乏人类典型偏见,适合评估语篇理解
本文比较了单语和多语大语言模型(涵盖不同规模)生成的数据与人类参与者在实验中提供的数据,研究已确立的语篇偏见。目标是建立一个基准,以话语偏见作为更通用话语理解能力的可靠代理。具体考察了隐含因果动词,心理学语言学研究发现人类在三个现象上存在偏见:(i) 指代关系建立(实验1),(ii) 连贯关系(实验2),(iii) 特定指称表达使用(实验3和4)。在指代关系方面,仅最大规模的单语模型(德语 Bloom 6.4B)表现出类人偏见。在连贯关系方面,无一模型表现出人类常见的解释偏见。在指称表达方面,所有模型均偏好用更简单的形式指代主语而非宾语,但未观察到显著偏见效应,与近期人类偏见研究结果不一致。
原文摘要 · Abstract (English)
In this paper, we compare data generated with mono- and multilingual LLMs spanning a range of model sizes with data provided by human participants in an experimental setting investigating well-established discourse biases. Beyond the comparison as such, we aim to develop a benchmark to assess the capabilities of LLMs with discourse biases as a robust proxy for more general discourse understanding capabilities. More specifically, we investigated Implicit Causality verbs, for which psycholinguistic research has found participants to display biases with regard to three phenomena:\ the establishment of (i) coreference relations (Experiment 1), (ii) coherence relations (Experiment 2), and (iii) the use of particular referring expressions (Experiments 3 and 4). With regard to coreference biases we found only the largest monolingual LLM (German Bloom 6.4B) to display more human-like biases. For coherence relation, no LLM displayed the explanation bias usually found for humans. For referring expressions, all LLMs displayed a preference for referring to subject arguments with simpler forms than to objects. However, no bias effect on referring expression was found, as opposed to recent studies investigating human biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。