测试大模型在量刑辅助中是否出现偏见,发现其对受害者美德和头衔光环有轻微偏好。
Assessing Cognitive Biases in LLMs for Judicial Decision Support: Virtuous Victim and Halo Effects
- 用改写过的案件描述隔离变量,测试模型对偏见的响应
- 模型对受害者美德效应反应更强,但对附加同意无显著惩罚
- 除学历光环外,其他头衔光环影响略低于人类,适合司法辅助参考
我们研究大型语言模型(LLMs)是否表现出类似人类的认知偏见,重点关注其在量刑决策——这一要求公平性的决策系统——中的潜在影响。选取两种关键偏见:受害者美德效应(VVE),重点考察相邻同意是否存在时该效应的减弱;以及基于声望的晕轮效应(职业、公司、资质)。使用从已有文献改编的案例片段,避免模型从训练数据中回忆,通过保持其他细节一致,仅改变特定因素,测量结果差异百分比。评估了五种代表性模型(ChatGPT 5 Instant、ChatGPT 5 Thinking、DeepSeek V3.1、Claude Sonnet 4、Gemini 2.5 Flash),进行独立多轮实验。研究发现,存在较大的受害者美德效应,但相邻同意并未带来统计显著的惩罚效应;与人类相比,声望相关的晕轮效应略有减弱,但以资质为基础的声望影响显著降低。尽管不同模型间结果差异限制了当前司法应用,整体表现仍优于人类基准。
原文摘要 · Abstract (English)
We investigate whether large language models (LLMs) display human-like cognitive biases, focusing on potential implications for assistance in judicial sentencing, a decision-making system where fairness is paramount. Two of the most relevant biases were chosen: the virtuous victim effect (VVE), with emphasis given to its reduction when adjacent consent is present, and prestige-based halo effects (occupation, company, and credentials). Using vignettes that were altered from prior literature to avoid LLMs recalling from their training data, we isolate each manipulation by holding all other details consistent, then measuring the percentage difference in outcomes. Five models were evaluated as representative LLMs in independent multi-run trials per condition (ChatGPT 5 Instant, ChatGPT 5 Thinking, DeepSeek V3.1, Claude Sonnet 4, Gemini 2.5 Flash). Our research discovers that there is larger VVE, there is no statistically significant penalty for adjacent-consent, and the halo effect is slightly reduced when compared to humans, with an exception for credential based prestige, which had a large reduction. Despite the variation across different models and outputs restricting current judicial usage, there were modest improvements compared to human benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。