arXiv:2504.03803cs.CLcs.CY2025-04被引 19

14个主流大模型在政治话题上存在审查,且多针对本地用户。

What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices

  • 区分硬性拒绝与软性删减,分析模型对政治人物的回应差异。
  • 跨国家模型审查倾向不同,多数仅采用硬或软一种方式。
  • 适合关注AI偏见、信息控制及多国模型对比的研究者。

大型语言模型(LLMs)日益成为信息入口,但其内容审核实践仍缺乏研究。本文考察了当被问及政治话题时,这些模型拒绝回答或省略信息的程度。我们区分了硬性审查(如生成拒绝、错误提示或预设否认回复)和软性审查(即选择性省略或淡化关键内容),并在涉及广泛政治人物的问题中识别出这些行为。分析涵盖来自西方国家、中国和俄罗斯的14个前沿模型,并使用联合国六种官方语言进行测试。结果显示,尽管所有模型均存在审查现象,但主要针对其本土受众,通常表现为硬性或软性审查之一(极少同时出现)。该发现凸显了公开可用模型在意识形态与地理多样性上的不足,以及提升模型审核策略透明度以支持用户知情选择的重要性。所有数据均免费开放。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed as gateways to information, yet their content moderation practices remain underexplored. This work investigates the extent to which LLMs refuse to answer or omit information when prompted on political topics. To do so, we distinguish between hard censorship (i.e., generated refusals, error messages, or canned denial responses) and soft censorship (i.e., selective omission or downplaying of key elements), which we identify in LLMs' responses when asked to provide information on a broad range of political figures. Our analysis covers 14 state-of-the-art models from Western countries, China, and Russia, prompted in all six official United Nations (UN) languages. Our analysis suggests that although censorship is observed across the board, it is predominantly tailored to an LLM provider's domestic audience and typically manifests as either hard censorship or soft censorship (though rarely both concurrently). These findings underscore the need for ideological and geographic diversity among publicly available LLMs, and greater transparency in LLM moderation strategies to facilitate informed user choices. All data are made freely available.

大模型审查政治敏感内容过滤多语言分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。