arXiv:2510.12217cs.CLcs.AI2025-10

为大模型公平性评估引入伤害严重度分级,更贴近真实部署场景。

HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment

  • 按伤害严重度将9个应用领域分为三档,构建部署对齐评估框架。
  • 8个大模型测试显示:模型规模不保证公平性,医学推理更优但教育表现差。
  • 适合关注模型实际落地风险的开发者与决策者参考。

大语言模型在医疗、法律、招聘和教育等高影响领域日益普及,部署前的公平性与偏见评估至关重要。然而,现有评估缺乏真实场景依据,且未区分伤害严重程度——例如手术决策中的偏见不应与文本摘要的风格偏见同等对待。为此,我们提出HALF(Harm-Aware LLM Fairness)框架,通过五阶段流程将九个应用领域划分为严重、中等、轻微三类伤害等级,实现面向部署的公平性评估。在八款大模型上的评估结果表明:(1) 模型在不同领域公平性表现不一致;(2) 模型规模或性能无法保证公平性;(3) 推理模型在医疗决策中表现更好,但在教育任务中表现更差。研究结论指出,现有基准测试的成功与实际部署准备之间存在显著差距。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed across high-impact domains, from clinical decision support and legal analysis to hiring and education, making fairness and bias evaluation before deployment critical. However, existing evaluations lack grounding in real-world scenarios and do not account for differences in harm severity, e.g., a biased decision in surgery should not be weighed the same as a stylistic bias in text summarization. To address this gap, we introduce HALF (Harm-Aware LLM Fairness), a deployment-aligned framework that assesses model bias in realistic applications and weighs the outcomes by harm severity. HALF organizes nine application domains into three tiers (Severe, Moderate, Mild) using a five-stage pipeline. Our evaluation results across eight LLMs show that (1) LLMs are not consistently fair across domains, (2) model size or performance do not guarantee fairness, and (3) reasoning models perform better in medical decision support but worse in education. We conclude that HALF exposes a clear gap between previous benchmarking success and deployment readiness.

大模型评估公平性部署对齐伤害分级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。