分析大模型对美国最高法院案件的政治倾向,发现其更像训练数据而非人类意见。
Better Aligned with Survey Respondents or Training Data? Unveiling Political Leanings of LLMs on U.S. Supreme Court Cases
- 通过量化方法评估预训练语料库的政治倾向
- 32个案件中模型倾向与训练数据强相关,与调查的人类观点无关
- 揭示了模型偏差源于数据记忆,提醒需加强数据审查
近期研究显示,大语言模型(LLMs)倾向于记忆和再现其训练数据中的模式与偏见,引发对其输出政治偏见的担忧。本文探究大模型的政治倾向是否反映其预训练语料库中的记忆模式。我们提出一种量化评估预训练语料库政治倾向的方法,并考察大模型的政治立场更倾向于训练数据还是调查中的公众意见。以32个美国最高法院案件为例,涵盖堕胎权、投票权等争议议题。结果显示,大模型的政治倾向与训练数据高度一致,但与人类调查意见无显著相关性。这表明模型偏见主要源自数据记忆,强调了负责任地筛选训练数据的重要性,并呼吁发展审计模型记忆化的技术以实现人机对齐。
原文摘要 · Abstract (English)
Recent works have shown that Large Language Models (LLMs) have a tendency to memorize patterns and biases present in their training data, raising important questions about how such memorized content influences model behavior. One such concern is the emergence of political bias in LLM outputs. In this paper, we investigate the extent to which LLMs' political leanings reflect memorized patterns from their pretraining corpora. We propose a method to quantitatively evaluate political leanings embedded in the large pretraining corpora. Subsequently we investigate to whom are the LLMs' political leanings more aligned with, their pretrainig corpora or the surveyed human opinions. As a case study, we focus on probing the political leanings of LLMs in 32 US Supreme Court cases, addressing contentious topics such as abortion and voting rights. Our findings reveal that LLMs strongly reflect the political leanings in their training data, and no strong correlation is observed with their alignment to human opinions as expressed in surveys. These results underscore the importance of responsible curation of training data, and the methodology for auditing the memorization in LLMs to ensure human-AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。