arXiv:2411.00585cs.CYcs.AI2024-11中稿 · ACM International …被引 23

测试大模型在角色扮演中是否产生偏见,发现普遍存在歧视性回答。

Fairness Testing of Large Language Models in Role-Playing

  • 用大模型生成550个社会角色和3.3万条问题,系统测试偏见
  • 10个主流大模型共检测出10.7万条偏见回应,单模型最高1.7万条
  • 数据集开源,适合安全评估与公平性研究者使用

大语言模型(LLMs)已成为现代语言驱动软件的核心,角色扮演是其关键应用之一。尽管已有研究揭示了模型输出中的社会偏见,但其在角色扮演场景中的表现仍不明确。本文开展了一项实证研究,通过大模型生成涵盖11种人口属性的550个社会角色,构建33,000条针对性问题,涵盖是/否、选择题与开放题,用于触发模型扮演特定角色并作答。采用规则与大模型结合的方法识别偏见,并经人工验证。对10个先进大模型的评估发现,共存在107,580条偏见回应,各模型间差异显著,最低7,579条,最高16,963条,凸显角色扮演中偏见的普遍性。为支持后续研究,数据集及所有脚本与结果已公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become foundational in modern language-driven software applications, profoundly influencing daily life. A critical technique in leveraging their potential is role-playing, where LLMs simulate diverse roles to enhance their real-world utility. However, while research has highlighted the presence of social biases in LLM outputs, it remains unclear whether and to what extent these biases emerge during role-playing scenarios. In this paper, we conduct an empirical study on fairness testing of LLMs in role-playing scenarios. To enable this testing, we use LLMs to generate 550 social roles spanning a comprehensive set of 11 demographic attributes, producing 33,000 role-specific questions that target various forms of bias. These questions, covering Yes/No, multiple-choice, and open-ended formats, are designed to prompt LLMs to adopt specific roles and respond accordingly. We employ a combination of rule-based and LLM-based strategies to identify biased responses, rigorously validated through human evaluation. Using the generated questions as the test cases, we conduct extensive evaluations of 10 advanced LLMs. The evaluation reveal 107,580 biased responses across the studied LLMs, with individual models yielding between 7,579 and 16,963 biased responses, underscoring the prevalence of bias in role-playing contexts. To support future research, we have publicly released the dataset, along with all scripts and experimental results.

大模型偏见角色扮演公平性测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。