arXiv:2503.05280cs.CL2025-03ACL被引 2

用NLP揭秘跨国内容审核背后的隐藏机制

Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing

  • 通过反向工程各国审核决策,构建跨国家文本分类器
  • 发现不同国家在不同时期的审查模式存在显著差异
  • 验证大模型生成解释的有效性,适合政策与算法研究者

自然语言处理(NLP)在多类别文本分类中的能力推动了其在在线内容审核(如仇恨言论、假新闻检测)中的应用。然而,目前对这些方法如何做出决策或为何特定内容被审核仍缺乏深入理解。为揭示内容审核的隐藏机制,本文探索三个方向:1)训练分类器以逆向解析多国内容审核决策;2)通过Shapley值与大模型引导的解释分析审核决策依据。研究聚焦于跨国审核决策,使用来自Twitter Stream Grab的预设语料库。实验揭示了被审查内容在国家间及时间上的有趣模式。通过对三种大模型生成解释的人工评估,检验了大模型在内容审核中的有效性。最后讨论了未来方向及本工作的局限性与伦理考量。代码与数据见https://github.com/causalNLP/censorship。

原文摘要 · Abstract (English)

The ability of Natural Language Processing (NLP) methods to categorize text into multiple classes has motivated their use in online content moderation tasks, such as hate speech and fake news detection. However, there is limited understanding of how or why these methods make such decisions, or why certain content is moderated in the first place. To investigate the hidden mechanisms behind content moderation, we explore multiple directions: 1) training classifiers to reverse-engineer content moderation decisions across countries; 2) explaining content moderation decisions by analyzing Shapley values and LLM-guided explanations. Our primary focus is on content moderation decisions made across countries, using pre-existing corpora sampled from the Twitter Stream Grab. Our experiments reveal interesting patterns in censored posts, both across countries and over time. Through human evaluations of LLM-generated explanations across three LLMs, we assess the effectiveness of using LLMs in content moderation. Finally, we discuss potential future directions, as well as the limitations and ethical considerations of this work. Our code and data are available at https://github.com/causalNLP/censorship

内容审核NLP跨国比较可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。