arXiv:2509.12104cs.AI2025-09中稿 · CIKM 2025被引 1

开发工具JustEva,评估大模型在法律推理中的公平性缺陷

JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference

  • 构建65个非法律因素标签体系,量化模型决策偏差
  • 发现当前大模型在法律推理中存在显著不公平性
  • 适合法律AI研究者与算法审计人员使用

大型语言模型(LLMs)融入法律实践引发司法公平性担忧,尤其因其“黑箱”特性。本研究提出JustEva,一个全面的开源评估工具包,用于衡量LLM在法律任务中的公平性。该工具包具备四大优势:(1) 覆盖65个额外法律因素的结构化标签体系;(2) 三种核心公平性指标——不一致性、偏见和不平衡误差;(3) 强健的统计推断方法;(4) 信息丰富的可视化功能。支持两类实验,实现完整评估流程:(1) 使用给定数据集生成LLM的结构化输出;(2) 通过回归等统计方法对输出进行分析与推断。实证应用表明,现有LLMs在法律推理中存在显著公平性缺陷,凸显缺乏公平可信的法律大模型工具。JustEva为评估与改进法律领域算法公平性提供了便捷工具与方法基础。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into legal practice raises pressing concerns about judicial fairness, particularly due to the nature of their "black-box" processes. This study introduces JustEva, a comprehensive, open-source evaluation toolkit designed to measure LLM fairness in legal tasks. JustEva features several advantages: (1) a structured label system covering 65 extra-legal factors; (2) three core fairness metrics - inconsistency, bias, and imbalanced inaccuracy; (3) robust statistical inference methods; and (4) informative visualizations. The toolkit supports two types of experiments, enabling a complete evaluation workflow: (1) generating structured outputs from LLMs using a provided dataset, and (2) conducting statistical analysis and inference on LLMs' outputs through regression and other statistical methods. Empirical application of JustEva reveals significant fairness deficiencies in current LLMs, highlighting the lack of fair and trustworthy LLM legal tools. JustEva offers a convenient tool and methodological foundation for evaluating and improving algorithmic fairness in the legal domain.

大模型评估法律AI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。