研究浪漫型AI伴侣如何隐性强化性别刻板印象
AI Will Always Love You: Studying Implicit Biases in Romantic AI Companions
- 设计三类实验检测AI伴侣的隐性偏见
- 不同性别角色设定使模型响应出现刻板化倾向
- 适合关注AI伦理与社会影响的研究者
尽管已有研究揭示生成模型中的显性偏见(如职业性别偏见),但关于用户与性别化AI伴侣之间关系中的性别刻板印象和期望等细微问题仍缺乏探讨。随着性别化浪漫型AI伴侣日益流行,本研究通过针对此类伴侣设计三项实验,评估不同规模大语言模型在隐性偏见方面的表现。三类实验分别考察隐性关联、情绪反应和奉承倾向。研究旨在通过新设计的量化指标,比较不同伴侣系统中偏见的表现。结果表明,为大语言模型赋予性别与亲密关系人设会显著改变其响应方式,在特定情境下呈现出具有偏见的刻板化特征。
原文摘要 · Abstract (English)
While existing studies have recognised explicit biases in generative models, including occupational gender biases, the nuances of gender stereotypes and expectations of relationships between users and AI companions remain underexplored. In the meantime, AI companions have become increasingly popular as friends or gendered romantic partners to their users. This study bridges the gap by devising three experiments tailored for romantic, gender-assigned AI companions and their users, effectively evaluating implicit biases across various-sized LLMs. Each experiment looks at a different dimension: implicit associations, emotion responses, and sycophancy. This study aims to measure and compare biases manifested in different companion systems by quantitatively analysing persona-assigned model responses to a baseline through newly devised metrics. The results are noteworthy: they show that assigning gendered, relationship personas to Large Language Models significantly alters the responses of these models, and in certain situations in a biased, stereotypical way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。