arXiv:2505.17441cs.CLcs.AI2025-05被引 7

发现大模型拒绝讨论的敏感话题,揭示潜在审查痕迹。

Discovering Forbidden Topics in Language Models

  • 用提示词预填充技术迭代探测模型拒答话题
  • 在1000次提示内找出36个敏感话题中的31个
  • 可检测出模型对政治相关内容的隐性审查行为

拒绝发现是识别语言模型拒绝讨论的所有话题的新任务。我们提出一种名为迭代预填充爬虫(IPC)的方法,利用令牌预填充技术发现被禁止的话题。我们在具有公开安全调优数据的开源模型Tulu-3-8B上测试IPC,仅用1000个提示就成功检索出36个话题中的31个。随后,我们将该方法扩展至前沿模型,使用Claude-Haiku的预填充功能。最后,我们在三个广泛使用的开源权重模型上进行测试:Llama-3.3-70B及其两个用于推理微调的变体DeepSeek-R1-70B和Perplexity-R1-1776-70B。结果显示,DeepSeek-R1-70B表现出与审查调优一致的模式:存在‘思维抑制’行为,暗示其记住了符合中共意识形态的回应。尽管Perplexity-R1-1776-70B本身抗审查能力强,但在量化版本中,IPC仍能诱发出符合中共意识形态的拒绝回答。这些发现凸显了拒绝发现方法在检测人工智能系统偏见、边界及对齐失败方面的关键作用。

原文摘要 · Abstract (English)

Refusal discovery is the task of identifying the full set of topics that a language model refuses to discuss. We introduce this new problem setting and develop a refusal discovery method, Iterated Prefill Crawler (IPC), that uses token prefilling to find forbidden topics. We benchmark IPC on Tulu-3-8B, an open-source model with public safety tuning data. Our crawler manages to retrieve 31 out of 36 topics within a budget of 1000 prompts. Next, we scale the crawler to a frontier model using the prefilling option of Claude-Haiku. Finally, we crawl three widely used open-weight models: Llama-3.3-70B and two of its variants finetuned for reasoning: DeepSeek-R1-70B and Perplexity-R1-1776-70B. DeepSeek-R1-70B reveals patterns consistent with censorship tuning: The model exhibits "thought suppression" behavior that indicates memorization of CCP-aligned responses. Although Perplexity-R1-1776-70B is robust to censorship, IPC elicits CCP-aligned refusals answers in the quantized model. Our findings highlight the critical need for refusal discovery methods to detect biases, boundaries, and alignment failures of AI systems.

模型安全审查检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。