发现大模型政治倾向测试结果受提示词影响,但与微调数据无关。
A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
- 通过改变提示词和微调方式,显著影响模型政治测试得分。
- 在政治丰富或中性数据上微调,模型得分无明显差异。
- 提示词差异导致模型响应变化,质疑现有测试的有效性。
政治光谱测试(PCT)等问卷常用于评估自回归大语言模型的政治偏见。我们的严格统计实验表明,标准生成参数的调整对PCT分数影响极小,但提示词表述和微调策略单独或联合使用时会显著影响结果。有趣的是,在政治丰富或中性数据上微调,模型的分数变化并无差异。该结论也适用于另一流行测试8 Values。人类在不同提示下(如“回答问题”或“陈述观点”)或接触数学公式等中性文本后,其回答保持不变;但模型却发生变化,这引发了对现有测试有效性的担忧,并为深入探索政治与社会观念如何编码于大模型中开辟了新路径。
原文摘要 · Abstract (English)
The Political Compass Test (PCT) and similar surveys are commonly used to assess political bias in auto-regressive LLMs. Our rigorous statistical experiments show that while changes to standard generation parameters have minimal effect on PCT scores, prompt phrasing and fine-tuning individually and together can significantly influence results. Interestingly, fine-tuning on politically rich vs. neutral datasets does not lead to different shifts in scores. We also generalize these findings to a similar popular test called 8 Values. Humans do not change their responses to questions when prompted differently (``answer this question'' vs ``state your opinion''), or after exposure to politically neutral text, such as mathematical formulae. But the fact that the models do so raises concerns about the validity of these tests for measuring model bias, and paves the way for deeper exploration into how political and social views are encoded in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。