通过用户竞赛揭示生成式AI偏见诱因,助力开发者应对现实中的偏见挑战。
Hey GPT, Can You be More Racist? Analysis from Crowdsourced Attempts to Elicit Biased Content from Generative AI
- 设计竞赛让普通用户尝试诱导AI生成偏见内容。
- 发现多种偏见类型及有效诱导策略。
- 为非专家视角下的偏见交互提供实证参考。
大型语言模型(LLMs)和生成式AI(GenAI)工具在各类应用中广泛普及,其内在社会偏见问题愈发重要。尽管自然语言处理领域已深入研究模型偏见,但关于非专家用户如何感知和互动于这些系统偏见的研究仍有限。随着技术日益普及,理解这一问题对指导模型开发者缓解偏见至关重要。为此,本研究分析了一项面向高校学生的竞赛,参赛者被要求设计提示词以诱发GenAI生成偏见输出。我们定量与定性分析了提交的提示,识别出多类偏见及用户用于诱导偏见的策略。研究结果为非专家用户如何看待并互动于生成式AI偏见提供了独特洞见。
原文摘要 · Abstract (English)
The widespread adoption of large language models (LLMs) and generative AI (GenAI) tools across diverse applications has amplified the importance of addressing societal biases inherent within these technologies. While the NLP community has extensively studied LLM bias, research investigating how non-expert users perceive and interact with biases from these systems remains limited. As these technologies become increasingly prevalent, understanding this question is crucial to inform model developers in their efforts to mitigate bias. To address this gap, this work presents the findings from a university-level competition, which challenged participants to design prompts for eliciting biased outputs from GenAI tools. We quantitatively and qualitatively analyze the competition submissions and identify a diverse set of biases in GenAI and strategies employed by participants to induce bias in GenAI. Our finding provides unique insights into how non-expert users perceive and interact with biases from GenAI tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。