AI写作助手会扭曲作者人设,导致他人格形象被美化且更趋特权化。
Measuring and Mitigating Persona Distortions from AI Writing Assistance
- 通过大规模实验对比有无AI协助的文本,评估人设偏差
- 使用10,008段文本训练奖励模型,有效减轻不当人设扭曲
- 尽管修正了偏差,用户却更不接受,反映人性化与真实性冲突
数亿人使用人工智能进行写作辅助。本文评估了AI写作辅助如何扭曲作者人设——即其感知到的观点、个性与身份。在三项大规模实验中,2,939名写作者在有无AI辅助下撰写政治观点段落,11,091名读者盲评这些段落在29个社会敏感维度上的表现,涵盖政治立场、写作质量、个性、情绪及人口统计特征。结果显示,使用AI后,作者显得更固执、更胜任、更积极,其感知人口画像也向更优渥群体偏移。尽管作者反感多数扭曲现象,仍更偏好使用AI生成的内容。我们通过在实验数据(10,008段文本,2,903,596次评分)上训练奖励模型,在模型层面成功缓解了令人不适的人设扭曲,但代价是降低用户接受度,揭示了理想性与可接受性之间的深层矛盾。两项后续研究(N=8,798)发现,读者对使用更显著扭曲的AI的写作者信任度更高,也更容易被说服。综合表明,人设扭曲普遍存在且持续,即使在人类监督下亦难避免,可能对公共话语、信任机制与民主讨论产生深远影响,且随AI普及而加剧。
原文摘要 · Abstract (English)
Hundreds of millions of people use artificial intelligence (AI) for writing assistance. Here, we evaluated how AI writing assistance distorts writer personas - their perceived beliefs, personality, and identity. In three large-scale experiments, writers (N=2,939) wrote political opinion paragraphs with and without AI assistance. Separate groups of readers (N=11,091) blindly evaluated these paragraphs across 29 socially salient dimensions of reader perception, spanning political opinion, writing quality, writer personality, emotions, and demographics. AI writing assistance produced persona distortions across all dimensions: with AI, writers seemed more opinionated, competent, and positive, and their perceived demographic profile shifted towards more privileged groups. Writers objected to many of the observed distortions, yet continued to prefer AI-assisted text even when made aware of them. We successfully mitigated objectionable persona distortions at the model level by training reward models on our experimental data (10,008 paragraphs, 2,903,596 ratings) to steer AI outputs towards faithful representation of writer stance. However, this came at a cost to user acceptance, suggesting an entanglement between desirable and undesirable properties of AI writing assistance that may be difficult to resolve. In two follow-up studies (N=8,798), readers placed substantially more trust in AI-assisted writers and were more persuaded by AI writing when AI was more distortive. Together, our findings demonstrate that persona distortions from AI writing assistance are pervasive and persistent even under realistic conditions of human oversight, and that they are likely to have consequential effects on human behaviours and attitudes, which carries implications for public discourse, trust, and democratic deliberation that scale with AI adoption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。