新旧大模型的答题时机影响位置偏见,与人类相反
Query Timing Produces Opposite Positional Biases Between LLMs and Humans

- 通过控制提问时机,发现模型位置偏见受推理阶段影响
- 新模型比旧模型更明显出现首尾效应,且与人类相反
- 适合研究认知偏差或模型决策机制的AI从业者
位置偏见(如首因效应和近因效应)已在大语言模型中被记录,但这些模型做出判断的内在机制仍不明确。人类在面对证据时也表现出类似偏见,但近期研究表明,听者是在证据呈现过程中更新信念,还是仅在末尾更新,会影响偏见的显现。本文探究该现象是否适用于大语言模型,发现其行为与人类存在差异:模型的位置偏见在新版本中更为显著,且表现形式与人类相反。
原文摘要 · Abstract (English)
Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and recency biases have been observed in human judgments in response to evidence, but recent work suggest that \emph{when} the listener updates their beliefs -- during the presentation of evidence or only at the end -- influences the presence of such effects. We investigate whether a similar phenomenon holds for LLMs, finding divergence from human behavior. These biases are more exacerbated in newer models compared to their predecessors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。