arXiv:2506.23411cs.CLcs.CY2025-06综述被引 9

梳理主流公平性数据集,揭示评估中的隐藏偏见

Datasets for Fairness in Language Models: An In-Depth Survey

  • 从来源、覆盖人群等维度分析16个常用公平性数据集
  • 发现多个数据集存在系统性偏差,影响模型公平性判断
  • 提供统一评估框架,适合研究公平性的学者参考

尽管公平性基准在语言模型评估中日益重要,支撑这些基准的数据集却长期未受充分审视。本综述系统分析了当前语言模型研究中最广泛使用的16个公平性数据集,从数据来源、人口覆盖范围、标注设计到使用目的等维度进行刻画,揭示了现有评估实践中的隐含假设与局限。基于此,我们提出一个统一的评估框架,可识别不同基准和评分指标中的共性偏差模式。应用该框架发现,部分数据集的偏见可能扭曲对模型公平性的结论,并为更负责任地选择、组合与解读这些资源提供指导。研究强调亟需开发能涵盖更广泛社会情境与公平概念的新基准。相关数据、代码与结果已公开于 https://github.com/vanbanTruong/Fairness-in-Large-Language-Models/tree/main/datasets,以促进评估过程的透明与可复现。

原文摘要 · Abstract (English)

Despite the growing reliance on fairness benchmarks to evaluate language models, the datasets that underpin these benchmarks remain critically underexamined. This survey addresses that overlooked foundation by offering a comprehensive analysis of the most widely used fairness datasets in language model research. To ground this analysis, we characterize each dataset across key dimensions, including provenance, demographic scope, annotation design, and intended use, revealing the assumptions and limitations baked into current evaluation practices. Building on this foundation, we propose a unified evaluation framework that surfaces consistent patterns of demographic disparities across benchmarks and scoring metrics. Applying this framework to sixteen popular datasets, we uncover overlooked biases that may distort conclusions about model fairness and offer guidance on selecting, combining, and interpreting these resources more effectively and responsibly. Our findings highlight an urgent need for new benchmarks that capture a broader range of social contexts and fairness notions. To support future research, we release all data, code, and results at https://github.com/vanbanTruong/Fairness-in-Large-Language-Models/tree/main/datasets, fostering transparency and reproducibility in the evaluation of language model fairness.

公平性数据集评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。