提出用多元理论分析大模型偏见,揭示其同质化危害并设计多样性增强方法。
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
- 从性别与酷儿理论出发,构建评估大模型同质化的框架
- 实验发现Claude 3.5 Haiku在开放式故事生成中存在显著性别刻板印象
- 提出'异质再生产'任务以主动促进多样性,适合关注AI伦理的研究者
生成式AI模型会复现并放大训练数据中的社会偏见,导致模式坍缩等机制引发同质化问题。这种多样性丧失不仅伤害边缘群体,也损害全体用户的体验。我们主张同质化应成为人工智能安全的核心议题。为此,提出一个可嵌入不同价值体系的框架,用于有意义地表征大语言模型中的同质化现象。通过一项实验,在开放式故事生成任务中揭示了Claude 3.5 Haiku模型中存在的性别偏见。基于酷儿理论,将同质化形式化为规范性问题;借鉴女性主义话语,引入‘异质再生产’作为一类缓解同质化的任务。本工作开启了一条协作研究路径,旨在理解并推进人工智能中的多样性。
原文摘要 · Abstract (English)
Generative AI models reproduce the human biases in their training data and further amplify them through mechanisms such as mode collapse. The loss of diversity produces homogenization, which not only harms the minoritized but impoverishes everyone. We argue homogenization should be a central concern in AI safety. To meaningfully characterize homogenization in Large Language Models (LLMs), we introduce a framework that allows stakeholders to encode their context and value system. We illustrate our approach with an experiment that surfaces gender bias in an LLM (Claude 3.5 Haiku) on an open-ended story prompt. Building from queer theory, we formalize homogenization in terms of normativity. Borrowing language from feminist theory, we introduce the concept of xeno-reproduction as a class of tasks for mitigating homogenization by promoting diversity. Our work opens a collaborative line of research that seeks to understand and advance diversity in AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。