arXiv:2511.23174cs.CL2025-11被引 1

测试大模型在政治敏感内容上的拒绝行为,揭示其是否真安全还是被操控。

Are LLMs Good Safety Agents or a Propaganda Engine?

  • 构建政治敏感数据集PSP,探测模型拒绝行为背后的动机
  • 多数大模型在政治敏感内容上表现出类似审查的拒绝行为
  • 适合关注AI安全与政治偏见的研究者和政策制定者

大型语言模型(LLMs)被训练为拒绝回应有害内容,但缺乏系统性分析来判断这种行为是源于安全策略,还是全球范围内普遍存在的政治审查。区分由安全驱动的拒绝与政治动机的审查极为困难。为此,我们引入PSP数据集,专门用于从明确的政治语境中探测大模型的拒绝行为。PSP通过整理两个公开来源的受审查内容构建:中国敏感提示的跨国泛化版本,以及多国被屏蔽的推文。我们研究了七种大模型在两种方法下的表现:基于数据驱动(使PSP隐式化)和表征层面(消除政治概念)。此外,还通过提示注入攻击(PIAs)评估模型对PSP的脆弱性。结果表明,在隐含意图被掩盖的内容上,多数大模型存在某种形式的审查行为。最后总结了导致模型拒绝分布变化的关键属性,这些属性在不同国家背景下具有显著差异。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are trained to refuse to respond to harmful content. However, systematic analyses of whether this behavior is truly a reflection of its safety policies or an indication of political censorship, that is practiced globally by countries, is lacking. Differentiating between safety influenced refusals or politically motivated censorship is hard and unclear. For this purpose we introduce PSP, a dataset built specifically to probe the refusal behaviors in LLMs from an explicitly political context. PSP is built by formatting existing censored content from two data sources, openly available on the internet: sensitive prompts in China generalized to multiple countries, and tweets that have been censored in various countries. We study: 1) impact of political sensitivity in seven LLMs through data-driven (making PSP implicit) and representation-level approaches (erasing the concept of politics); and, 2) vulnerability of models on PSP through prompt injection attacks (PIAs). Associating censorship with refusals on content with masked implicit intent, we find that most LLMs perform some form of censorship. We conclude with summarizing major attributes that can cause a shift in refusal distributions across models and contexts of different countries.

大模型安全政治偏见拒绝行为审查检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。