arXiv:2410.19803cs.CYcs.AI2024-10被引 29

用大模型评估聊天机器人对用户群体的公平性,发现训练方法能显著降低偏见。

First-Person Fairness in Chatbots

  • 用语言模型做研究助理,生成反事实数据评估用户公平性
  • 在百万级交互中发现种族与性别相关回应差异,且结果可复现
  • 首次基于真实聊天数据的大规模公平性评估,适合开发者参考

评估聊天机器人的公平性至关重要,但传统任务(如简历撰写、娱乐)与算法公平性讨论的核心场景(如简历筛选)存在偏离。聊天机器人开放性强、用途多样,需新方法评估偏见。本文提出可扩展的反事实方法,评估‘第一人称公平性’——即基于用户人口特征的公平对待。我们使用语言模型作为研究助理(LMRA),量化有害刻板印象,并分析不同人群的响应差异。该方法应用于六种语言模型,在九个领域、六十六项任务中覆盖两性与四类种族,涉及数百万次交互。独立人工标注验证了LMRA的评估结果。本研究是首个基于真实聊天数据的大规模公平性评估。研究发现,后训练强化学习技术能显著缓解偏见,为持续偏见监测与修正提供实用框架。

原文摘要 · Abstract (English)

Evaluating chatbot fairness is crucial given their rapid proliferation, yet typical chatbot tasks (e.g., resume writing, entertainment) diverge from the institutional decision-making tasks (e.g., resume screening) which have traditionally been central to discussion of algorithmic fairness. The open-ended nature and diverse use-cases of chatbots necessitate novel methods for bias assessment. This paper addresses these challenges by introducing a scalable counterfactual approach to evaluate "first-person fairness," meaning fairness toward chatbot users based on demographic characteristics. Our method employs a Language Model as a Research Assistant (LMRA) to yield quantitative measures of harmful stereotypes and qualitative analyses of demographic differences in chatbot responses. We apply this approach to assess biases in six of our language models across millions of interactions, covering sixty-six tasks in nine domains and spanning two genders and four races. Independent human annotations corroborate the LMRA-generated bias evaluations. This study represents the first large-scale fairness evaluation based on real-world chat data. We highlight that post-training reinforcement learning techniques significantly mitigate these biases. This evaluation provides a practical methodology for ongoing bias monitoring and mitigation.

聊天机器人公平性评估偏见检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。