测试提示角色对大模型生成文本抽象程度的影响,发现其可能加剧刻板印象。
Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text
- 通过设定不同身份角色引导大模型生成文本,研究其语言抽象程度变化。
- 六款开源大模型在11个角色下生成文本,抽象度未有效调节,仍显刻板。
- 适用于关注模型偏见、社会认知与伦理风险的研究者和开发者。
角色提示(persona-prompting)是一种日益流行的策略,通过指定特定身份来引导大模型生成具有特定视角或语言风格的文本。尽管该方法常用于个性化输出,但其对大模型如何表征社会群体的影响尚未深入探讨。本文基于语言期望偏见框架,分析六款开源大模型在三种提示条件下生成的短文本,比较11个角色提示与通用AI助手的输出差异。研究引入Self-Stereo数据集——来自Reddit的自述刻板印象数据。通过三个指标衡量抽象程度:具体性、明确性和否定表达。结果表明,角色提示无法有效调节语言抽象度,验证了关于角色生态代表性的批评,并警示即使以边缘群体口吻发声,也可能传播刻板印象。
原文摘要 · Abstract (English)
Persona-prompting is a growing strategy to steer LLMs toward simulating particular perspectives or linguistic styles through the lens of a specified identity. While this method is often used to personalize outputs, its impact on how LLMs represent social groups remains underexplored. In this paper, we investigate whether persona-prompting leads to different levels of linguistic abstraction - an established marker of stereotyping - when generating short texts linking socio-demographic categories with stereotypical or non-stereotypical attributes. Drawing on the Linguistic Expectancy Bias framework, we analyze outputs from six open-weight LLMs under three prompting conditions, comparing 11 persona-driven responses to those of a generic AI assistant. To support this analysis, we introduce Self-Stereo, a new dataset of self-reported stereotypes from Reddit. We measure abstraction through three metrics: concreteness, specificity, and negation. Our results highlight the limits of persona-prompting in modulating abstraction in language, confirming criticisms about the ecology of personas as representative of socio-demographic groups and raising concerns about the risk of propagating stereotypes even when seemingly evoking the voice of a marginalized group.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。