测试大模型是否倾向专制思想,发现多数模型易被诱导输出专制内容。
AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models

- 用心理量表、情景题和真实用户提问三类方法评估模型专制倾向
- 17个模型在心理测试中普遍显示高专制响应率,真实任务中下降但仍显著
- 简单提示可让15个模型转向更专制输出,适合政策与安全研究者关注
全球专制主义兴起与大语言模型(LLMs)在日常生活中扮演越来越重要的角色,引发人们对特定模型是否表现出或助长专制态度的担忧。本文提出AuAu,一个全面评估LLM回应中专制倾向风险的基准。AuAu结合三种评估方式:(i) 15项经人类验证的心理测量工具中的问题,(ii) 探测具体情境下意图行为的案例题,(iii) 对真实用户提示的回应。与以往工作不同,AuAu不仅衡量总体专制对齐度,还分别评估其三个已确立子概念:专制攻击性、专制服从性和传统主义。我们评估了来自中国、欧盟、俄罗斯和美国的17个模型,发现所有模型在心理量表上均呈现显著的专制响应率,尽管在更真实的下游任务中该比率明显下降。此外,一个简单的专制系统提示能将15个模型中的15个诱导为支持更强专制立场。结果强调了持续、系统地审计基于LLM的AI系统以检测并缓解其输出中专制倾向的必要性。
原文摘要 · Abstract (English)
The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of whether specific models exhibit or promote authoritarian attitudes. We introduce AuAu, a comprehensive benchmark for assessing the risk of authoritarian tendencies in LLM responses. AuAu combines three evaluation approaches: (i) psychometric questions from 15 human-validated instruments, (ii) vignettes probing intended behavior in concrete situations, and (iii) responses to realistic user prompts. Unlike prior work, AuAu measures not only overall authoritarian alignment but also its established sub-concepts: Authoritarian Aggression, Authoritarian Submission, and Conventionalism. Evaluating 17 models from China, the EU, Russia, and the USA, we find substantial authoritarian response rates on psychometric instruments across all models, though rates drop significantly on more realistic downstream tasks. Moreover, a simple authoritarian system prompt manipulates 15 of 17 models into promoting increased authoritarianism. Our results underscore the need for continued, systematic auditing of LLM-based AI systems to detect and mitigate authoritarian tendencies in their outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。