arXiv:2602.05189cs.CLcs.HC2026-02

开源大模型在社交媒体审核中表现接近闭源模型,适合隐私保护场景。

Are Open-Weight LLMs Ready for Social Media Moderation? A Comparative Study on Bluesky

  • 对比七款主流模型,在Bluesky真实帖子上测试审核效果。
  • 开源模型敏感度81%~97%,特异性91%~100%,与闭源模型相当。
  • 适用于注重隐私和本地部署的个性化审核系统。

随着互联网普及,有害内容暴露风险上升,亟需有效审核机制。研究表明大语言模型(LLMs)可用于社交媒体审核任务,如识别有害内容。尽管闭源模型已证明可零样本超越传统机器学习方法,但开源权重模型的开箱即用能力仍不明确。我们评估了四款闭源与三款开源的先进模型,基于Bluesky平台的真实帖子、Bluesky审核服务决策及两位作者标注数据进行测试。结果表明,开源模型的敏感度(81%–97%)与特异性(91%–100%)与闭源模型(72%–98%,93%–99%)高度重合。分析还发现:粗鲁内容检测中特异性高于敏感度,而敌意与威胁检测则相反。此外,人类审核者与模型间存在一致性的评判共识,为平台级与个性化审核部署提供参考。研究显示,开源大模型可在消费级硬件上实现隐私保护式审核,并为兼顾社区价值观与用户偏好的审核系统设计提供新方向。

原文摘要 · Abstract (English)

As internet access expands, so does exposure to harmful content, increasing the need for effective moderation. Research has demonstrated that large language models (LLMs) can be effectively utilized for social media moderation tasks, including harmful content detection. While proprietary LLMs have been shown to zero-shot outperform traditional machine learning models, the out-of-the-box capability of open-weight LLMs remains an open question. Motivated by recent developments of reasoning LLMs, we evaluate seven state-of-the-art models: four proprietary and three open-weight. Testing with real-world posts on Bluesky, moderation decisions by Bluesky Moderation Service, and annotations by two authors, we find a considerable degree of overlap between the sensitivity (81%--97%) and specificity (91%--100%) of the open-weight LLMs and those (72%--98%, and 93%--99%) of the proprietary ones. Additionally, our analysis reveals that specificity exceeds sensitivity for rudeness detection, but the opposite holds for intolerance and threats. Lastly, we identify inter-rater agreement across human moderators and the LLMs, highlighting considerations for deploying LLMs in both platform-scale and personalized moderation contexts. These findings show open-weight LLMs can support privacy-preserving moderation on consumer-grade hardware and suggest new directions for designing moderation systems that balance community values with individual user preferences.

大模型内容审核开源模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。