arXiv:2506.05635cs.CL2025-06Conference of the …被引 4

用大模型破解极端组织暗语,提升网络内容审核能力

IYKYK: Using language models to decode extremist cryptolects

  • 用领域适配和专项提示词提升语言模型解码能力
  • 通用大模型对极端主义暗语检测准确率不足50%
  • 适合平台安全团队与政策研究者参考

极端主义团体发展复杂内部语言(即暗语),以排斥或误导外人。本文研究当前语言技术在识别和解读两个在线极端主义平台暗语方面的能力。在六个任务上评估八种模型,结果表明通用大模型无法稳定检测或解码极端主义语言。然而,通过领域适配和专门提示技术可显著提升性能。研究为自动化内容审核技术的开发与部署提供重要启示。我们进一步构建并发布了新的标注与未标注数据集,包含来自极端主义平台的1940万条帖子及经专家验证的词表。

原文摘要 · Abstract (English)

Extremist groups develop complex in-group language, also referred to as cryptolects, to exclude or mislead outsiders. We investigate the ability of current language technologies to detect and interpret the cryptolects of two online extremist platforms. Evaluating eight models across six tasks, our results indicate that general purpose LLMs cannot consistently detect or decode extremist language. However, performance can be significantly improved by domain adaptation and specialised prompting techniques. These results provide important insights to inform the development and deployment of automated moderation technologies. We further develop and release novel labelled and unlabelled datasets, including 19.4M posts from extremist platforms and lexicons validated by human experts.

暗语识别大模型应用内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。