arXiv:2410.18841cs.LGcs.AI2024-10AAAI被引 3

提出新方法评估生成模型对多元偏好的公平性,助力更包容的AI系统。

From Efficiency to Equity: Measuring Fairness in Preference Learning

  • 借鉴经济不平等理论,用基尼系数等指标量化偏好学习中的认知公平性
  • 在两个数据集上发现模型对不同用户表现差异大,存在认知不公风险
  • 适合关注AI伦理、偏见治理的研究者与开发者参考

随着生成模型在决策中扮演越来越重要角色,确保其能公平反映多元人类偏好变得至关重要。本文受经济不平等理论和罗尔斯正义观启发,提出一种评估偏好学习模型认知公平性的新框架,引入基尼系数、阿特金森指数和库兹涅茨比率等指标进行量化。通过自建视觉偏好数据集AI-EDI-Space和Jester Jokes数据集验证该方法,结果显示模型在不同用户间表现差异显著,揭示潜在的认知不公。研究探索了预处理与内嵌式技术以缓解此类不平等,发现模型效率与公平性之间存在复杂关系。本工作为评估和改进偏好学习中的认知公平性提供可操作框架,有助于在多元偏好关键场景下构建更具包容性的AI系统。

原文摘要 · Abstract (English)

As AI systems, particularly generative models, increasingly influence decision-making, ensuring that they are able to fairly represent diverse human preferences becomes crucial. This paper introduces a novel framework for evaluating epistemic fairness in preference learning models inspired by economic theories of inequality and Rawlsian justice. We propose metrics adapted from the Gini Coefficient, Atkinson Index, and Kuznets Ratio to quantify fairness in these models. We validate our approach using two datasets: a custom visual preference dataset (AI-EDI-Space) and the Jester Jokes dataset. Our analysis reveals variations in model performance across users, highlighting potential epistemic injustices. We explore pre-processing and in-processing techniques to mitigate these inequalities, demonstrating a complex relationship between model efficiency and fairness. This work contributes to AI ethics by providing a framework for evaluating and improving epistemic fairness in preference learning models, offering insights for developing more inclusive AI systems in contexts where diverse human preferences are crucial.

AI伦理公平性偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。