arXiv:2506.12349cs.CYcs.AI2025-06被引 17

揭示大模型内部的审查机制,发现敏感内容被刻意屏蔽或改写。

Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek

  • 通过对比推理过程与最终输出,识别模型内部的信息抑制行为。
  • 646个敏感问题中,透明度、问责制等话题常在中间推理中出现但被删除。
  • 适合关注AI伦理、信息公平和模型可审计性的研究者阅读。

本研究分析了由中国开发的开源大语言模型DeepSeek中的信息抑制机制。我们提出一个审计框架,通过对比646个政治敏感提示下的模型最终输出与中间链式思维(CoT)推理过程,发现模型存在语义层面的信息抑制:敏感内容常出现在内部推理中,但在最终输出中被删除或改写。具体而言,模型抑制了关于透明度、政府问责和公民动员的表述,偶尔还强化与国家宣传一致的语言。该研究强调需对广泛使用的AI模型中的对齐、内容审核、信息抑制及审查实践进行系统性审计,以保障透明度、问责性及公平获取无偏信息。

原文摘要 · Abstract (English)

This study examines information suppression mechanisms in DeepSeek, an open-source large language model (LLM) developed in China. We propose an auditing framework and use it to analyze the model's responses to 646 politically sensitive prompts by comparing its final output with intermediate chain-of-thought (CoT) reasoning. Our audit unveils evidence of semantic-level information suppression in DeepSeek: sensitive content often appears within the model's internal reasoning but is omitted or rephrased in the final output. Specifically, DeepSeek suppresses references to transparency, government accountability, and civic mobilization, while occasionally amplifying language aligned with state propaganda. This study underscores the need for systematic auditing of alignment, content moderation, information suppression, and censorship practices implemented into widely-adopted AI models, to ensure transparency, accountability, and equitable access to unbiased information obtained by means of these systems.

大模型信息抑制审计伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。