通过长期监测,揭示大模型对社会议题的拒绝策略变化。
Longitudinal Monitoring of LLM Content Moderation of Social Issues
- 构建系统持续追踪大模型对400+社会议题的拒绝行为。
- 发现未公开政策调整仍可被检测到,且不同公司/模型有差异。
- 适合关注AI治理、内容安全与模型透明度的研究者。
大型语言模型的输出受制于不透明且频繁变动的企业内容审核政策。模型对特定话题的拒绝生成不仅反映企业政策,也悄然影响公共话语。本文提出AI Watchman——一个纵向审计系统,用于公开测量和追踪大模型随时间的拒绝行为,以增强这一关键但黑箱化的环节的透明度。基于涵盖400多个社会议题的数据集,我们审计了OpenAI的GPT-4.1与GPT-5,以及DeepSeek(中英文版本)。结果表明,即使未公开的政策变更也可被AI Watchman检测到,并识别出公司与模型间的审核差异。我们还定性分析并分类了多种拒绝形式。本研究为大模型的纵向审计提供了实证支持,展示了AI Watchman作为实现该目标的可行系统。
原文摘要 · Abstract (English)
Large language models' (LLMs') outputs are shaped by opaque and frequently-changing company content moderation policies and practices. LLM moderation often takes the form of refusal; models' refusal to produce text about certain topics both reflects company policy and subtly shapes public discourse. We introduce AI Watchman, a longitudinal auditing system to publicly measure and track LLM refusals over time, to provide transparency into an important and black-box aspect of LLMs. Using a dataset of over 400 social issues, we audit Open AI's moderation endpoint, GPT-4.1, and GPT-5, and DeepSeek (both in English and Chinese). We find evidence that changes in company policies, even those not publicly announced, can be detected by AI Watchman, and identify company- and model-specific differences in content moderation. We also qualitatively analyze and categorize different forms of refusal. This work contributes evidence for the value of longitudinal auditing of LLMs, and AI Watchman, one system for doing so.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。