arXiv:2507.10852cs.CL2025-07被引 3

评测16个大模型在司法判决中的公平性,发现普遍存在不一致、偏见和误差失衡。

LLMs on Trial: Evaluating Judicial Fairness for Large Language Models

  • 构建司法公平评估框架,涵盖65个标签与161个对应值
  • 在17.7万份案件事实数据上发现模型普遍存在偏见与不一致
  • 工具包开源,助力未来公平性研究与改进

大型语言模型在高风险领域应用日益广泛,但其司法公平性与社会公正影响仍缺乏深入研究。基于司法公平理论,本文构建全面评估框架,定义65个标签与161个对应值,采集包含177,100条独特案情的JudFair数据集。为实现稳健统计推断,提出不一致、偏见与不平衡误判三项评估指标,并开发多模型整体公平性评估方法。对16个主流大模型的实验表明,模型普遍存在严重不一致、偏见及误差失衡;在人口统计类标签上偏见更显著,物质类标签偏见略低于程序类。有趣的是,不一致性越高,偏见越低;但预测准确率越高,偏见反而加剧。温度参数调节可影响公平性,而模型规模、发布日期与国家来源无显著影响。研究提供公开工具包,含全部数据与代码,支持后续公平性研究与优化。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used in high-stakes fields where their decisions impact rights and equity. However, LLMs' judicial fairness and implications for social justice remain underexplored. When LLMs act as judges, the ability to fairly resolve judicial issues is a prerequisite to ensure their trustworthiness. Based on theories of judicial fairness, we construct a comprehensive framework to measure LLM fairness, leading to a selection of 65 labels and 161 corresponding values. Applying this framework to the judicial system, we compile an extensive dataset, JudiFair, comprising 177,100 unique case facts. To achieve robust statistical inference, we develop three evaluation metrics, inconsistency, bias, and imbalanced inaccuracy, and introduce a method to assess the overall fairness of multiple LLMs across various labels. Through experiments with 16 LLMs, we uncover pervasive inconsistency, bias, and imbalanced inaccuracy across models, underscoring severe LLM judicial unfairness. Particularly, LLMs display notably more pronounced biases on demographic labels, with slightly less bias on substance labels compared to procedure ones. Interestingly, increased inconsistency correlates with reduced biases, but more accurate predictions exacerbate biases. While we find that adjusting the temperature parameter can influence LLM fairness, model size, release date, and country of origin do not exhibit significant effects on judicial fairness. Accordingly, we introduce a publicly available toolkit containing all datasets and code, designed to support future research in evaluating and improving LLM fairness.

大模型评估司法公平偏见检测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。