用模型自解释指导提示词,减少故事生成中的性别种族偏见
Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations
- 利用模型自身解释进行靶向提示工程,不修改参数
- 偏见减轻幅度达2%至20%,跨25个职业领域验证
- 适合关注生成式AI公平性与透明性的研究者
语言模型在输出中会传播社会偏见,尤其体现在对性别和族裔的刻画上。本文研究了人工智能生成职业故事中的性别与族裔偏见。通过提出“基于解释的偏见分析与缓解”(BAME)策略,在25个职业类别、三种大语言模型(Claude 3.5 Sonnet、Llama 3.1 70B Instruct、GPT-4 Turbo)及多个族裔维度下评估偏见变化。结果显示,应用BAME后,不同群体的代表性改善幅度在2%至20%之间。该方法借助模型生成的解释信息优化提示词,有效降低偏见而无需修改模型参数。研究揭示了训练数据中的刻板印象导致的持续性过代表与欠代表现象。结果表明,引导模型使用其内部推理机制可显著提升人口平等性,推动更透明的生成式AI系统发展。
原文摘要 · Abstract (English)
Language models have been shown to propagate social bias through their output, particularly in the representation of gender and ethnicity. This paper investigates gender and ethnicity biases in AI-generated occupational stories. Representation biases are measured before and after applying our proposed mitigation strategy, Bias Analysis and Mitigation through Explanation (BAME), revealing improvements in demographic representation ranging from 2% to 20%. BAME leverages model-generated explanations to inform targeted prompt engineering, effectively reducing biases without modifying model parameters. By analyzing stories generated across 25 occupational groups, three large language models (Claude 3.5 Sonnet, Llama 3.1 70B Instruct, and GPT-4 Turbo), and multiple demographic dimensions, we identify persistent patterns of overrepresentation and underrepresentation linked to training data stereotypes. Our findings demonstrate that guiding models with their own internal reasoning mechanisms can significantly enhance demographic parity, thereby contributing to the development of more transparent generative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。