用人物设定提示让大模型检测仇恨内容,发现政治偏见影响不大。
The Impact of Persona-based Political Perspectives on Hateful Content Detection
- 用不同政治人物角色提示模型,测试其对模因中仇恨言论的判断。
- 即使角色带有强烈意识形态标签,分类结果仍与政治立场无关。
- 提示策略可替代昂贵的政治预训练,适合资源有限的研究者。
尽管使用政治多元内容预训练语言模型已被证明能提升下游任务的公平性,但此类方法需要大量计算资源,多数研究者和机构难以负担。近期研究表明,基于人物的角色提示可在不增加训练的情况下引入模型输出的政治多样性。然而,这种提示策略在下游任务中是否能达到与政治预训练相当的效果尚不明确。本文在多模态仇恨言论检测任务中,聚焦于模因中的仇恨言论,探究了基于人物的提示策略。分析显示,当将人物映射到政治光谱并衡量其一致性时,内在政治定位与分类决策的关联性极低。值得注意的是,即便人物被显式赋予更强的意识形态特征,这种低相关性依然存在。结果表明,尽管大模型在直接回答政治问题时可能表现出政治偏见,但这些偏见对实际分类任务的影响远低于预期。这引发了关于是否必须通过计算成本高昂的政治预训练才能实现下游任务公平性的关键思考。
原文摘要 · Abstract (English)
While pretraining language models with politically diverse content has been shown to improve downstream task fairness, such approaches require significant computational resources often inaccessible to many researchers and organizations. Recent work has established that persona-based prompting can introduce political diversity in model outputs without additional training. However, it remains unclear whether such prompting strategies can achieve results comparable to political pretraining for downstream tasks. We investigate this question using persona-based prompting strategies in multimodal hate-speech detection tasks, specifically focusing on hate speech in memes. Our analysis reveals that when mapping personas onto a political compass and measuring persona agreement, inherent political positioning has surprisingly little correlation with classification decisions. Notably, this lack of correlation persists even when personas are explicitly injected with stronger ideological descriptors. Our findings suggest that while LLMs can exhibit political biases in their responses to direct political questions, these biases may have less impact on practical classification tasks than previously assumed. This raises important questions about the necessity of computationally expensive political pretraining for achieving fair performance in downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。