LLM做选择时会受顺序影响,好东西反而排后面被忽略。
Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models
- 用对比实验发现模型偏好受选项顺序影响,高质量时偏第一,低质量时偏最后。
- 在简历和选色任务中均观察到明显顺序偏差,且存在未被记录的姓名偏好。
- 提出新框架区分真实偏好与表面选择,适合用于高风险决策系统的评估。
大型语言模型(LLMs)正越来越多地应用于招聘、高校录取等高风险决策场景,这些场景常需在多个备选方案间做出选择。尽管已有研究指出LLM在比较中存在位置偏差,但这类偏差尚未得到系统分析,也未与底层偏好结构建立联系。本文首次全面研究了多种LLM在两个不同领域中的位置偏差:简历比较(代表真实高风险场景)和颜色选择(通过去除干扰因素隔离位置效应)。研究发现,存在强烈且一致的顺序效应,且呈现质量依赖特征——当所有选项质量较高时,模型更倾向于首选项;而当质量较低时,则偏好靠后的选项。此外,还发现一种此前未被记录的名称偏好,即某些名字即使控制了人口统计信号仍被偏爱。为区分表面的随机选择与真实的判断扭曲,本文扩展理性选择框架,将成对偏好分为稳健、脆弱或无差异三类。基于该框架,证明顺序效应可导致模型选择严格劣质的选项。结果表明,LLMs表现出人类决策中未见的独特失效模式。同时提出针对性缓解策略,包括创新性使用温度参数以恢复原始偏好。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in decision-support systems for high-stakes domains such as hiring and university admissions, where choices often involve selecting among competing alternatives. While prior work has noted position biases in LLM-driven comparisons, these biases have not been systematically analyzed or linked to underlying preference structures. We present the first comprehensive study of position biases across multiple LLMs and two distinct domains: resume comparisons, representing a realistic high-stakes context, and color selection, which isolates position effects by removing confounding factors. We find strong and consistent order effects, including a quality-dependent shift: when all options are high quality, models favor the first option, but when quality is lower, they favor later options. We also identify a previously undocumented bias: a name bias, where certain names are favored despite controlling for demographic signals. To separate superficial tie-breaking from genuine distortions of judgment, we extend the rational choice framework to classify pairwise preferences as robust, fragile, or indifferent. Using this framework, we show that order effects can lead models to select strictly inferior options. These results indicate that LLMs exhibit distinct failure modes not documented in human decision-making. We also propose targeted mitigation strategies, including a novel use of the temperature parameter, to recover underlying preferences when order effects distort model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。