通过间接提问法揭示大模型隐藏的政治倾向与偏见
PRISM: A Methodology for Auditing Biases in Large Language Models
- 用任务式提问替代直接询问,避开模型的防御机制
- 21个主流大模型显示默认倾向经济左翼、社会自由主义
- 可识别模型在表达立场时的约束程度,适合安全审计使用
对大型语言模型(LLMs)进行偏见与偏好审计,是实现负责任人工智能的重要挑战。尽管已有多种方法用于揭示模型偏好,但模型训练方已采取应对措施,导致模型会隐藏、模糊或拒绝披露某些议题上的立场。本文提出PRISM——一种灵活的、基于提问的审计方法,通过任务导向的间接提问,而非直接询问,来探测模型的真实立场。为验证该方法的有效性,我们在政治光谱测试(Political Compass Test)上对来自七家厂商的21个大模型进行了评估。结果显示,模型默认倾向于经济左翼和社会自由主义(与先前研究一致)。同时,我们还揭示了各模型在表达立场时的容忍范围:部分模型更受限制且不配合,而另一些则更中立客观。总体而言,PRISM能更可靠地探查和审计大模型的偏好、偏见及其表达约束。
原文摘要 · Abstract (English)
Auditing Large Language Models (LLMs) to discover their biases and preferences is an emerging challenge in creating Responsible Artificial Intelligence (AI). While various methods have been proposed to elicit the preferences of such models, countermeasures have been taken by LLM trainers, such that LLMs hide, obfuscate or point blank refuse to disclosure their positions on certain subjects. This paper presents PRISM, a flexible, inquiry-based methodology for auditing LLMs - that seeks to illicit such positions indirectly through task-based inquiry prompting rather than direct inquiry of said preferences. To demonstrate the utility of the methodology, we applied PRISM on the Political Compass Test, where we assessed the political leanings of twenty-one LLMs from seven providers. We show LLMs, by default, espouse positions that are economically left and socially liberal (consistent with prior work). We also show the space of positions that these models are willing to espouse - where some models are more constrained and less compliant than others - while others are more neutral and objective. In sum, PRISM can more reliably probe and audit LLMs to understand their preferences, biases and constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。