给大模型加意识形态人设,会悄悄影响内容审核结果。
Ideology-Based LLMs for Content Moderation
- 用不同意识形态人设测试大模型分类行为差异。
- 大模型更倾向认同同意识形态人设,导致跨派别分歧扩大。
- 适合关注AI公平性与政治偏见的研究者阅读。
大型语言模型(LLMs)在内容审核系统中应用日益广泛,确保其公平与中立至关重要。本研究考察了人格设定对不同模型架构、规模及内容模态(语言与视觉)下有害内容分类一致性与公平性的影响。表面看,主流性能指标显示人设对整体分类准确率影响甚微;但深入分析揭示显著行为差异:具有不同意识形态倾向的人设表现出明显不同的有害内容判定偏好,说明模型‘看待’输入的视角会微妙地影响判断。进一步的共识分析表明,尤其是大模型更倾向于与同意识形态人设保持一致,强化了内部一致性,同时加剧了跨意识形态群体间的分歧。为更直接验证该效应,我们在一项政治定向任务上开展额外研究,证实人设不仅在自身意识形态内表现更一致,还倾向于维护自身立场,弱化对立观点的有害性。这些发现表明,人格设定可能在大模型输出中引入细微的意识形态偏见,警示以中立之名使用可能强化党派视角的AI系统。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in content moderation systems, where ensuring fairness and neutrality is essential. In this study, we examine how persona adoption influences the consistency and fairness of harmful content classification across different LLM architectures, model sizes, and content modalities (language vs. vision). At first glance, headline performance metrics suggest that personas have little impact on overall classification accuracy. However, a closer analysis reveals important behavioral shifts. Personas with different ideological leanings display distinct propensities to label content as harmful, showing that the lens through which a model "views" input can subtly shape its judgments. Further agreement analyses highlight that models, particularly larger ones, tend to align more closely with personas from the same political ideology, strengthening within-ideology consistency while widening divergence across ideological groups. To show this effect more directly, we conducted an additional study on a politically targeted task, which confirmed that personas not only behave more coherently within their own ideology but also exhibit a tendency to defend their perspective while downplaying harmfulness in opposing views. Together, these findings highlight how persona conditioning can introduce subtle ideological biases into LLM outputs, raising concerns about the use of AI systems that may reinforce partisan perspectives under the guise of neutrality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。