arXiv:2505.20963cs.CLcs.AI2025-05被引 1

为德语报纸评论区设计上下文感知的自动内容审核模型

Context-Aware Content Moderation for German Newspaper Comments

  • 融合用户历史与文章主题的上下文信息提升分类效果
  • CNN与LSTM模型在上下文帮助下表现优于主流方法
  • 大模型零样本推理不依赖上下文且效果更差

在线讨论量持续增长,亟需高效的内容审核机制以维护负责任的网络对话。尽管社交媒体上的仇恨言论检测已较为成熟,针对德语报纸论坛的研究仍较有限。现有工作常忽略平台特有的上下文,如用户历史与文章主题。本文提出并评估了用于德语报纸论坛的二分类自动内容审核模型,引入上下文信息。基于奥地利《标准报》的百万条评论语料库(One Million Posts Corpus),采用LSTM、CNN及ChatGPT-3.5 Turbo进行实验。结果显示,CNN与LSTM模型能有效利用上下文信息,在性能上媲美当前最优方法;而ChatGPT的零样本分类在加入上下文后无提升,且表现较差。

原文摘要 · Abstract (English)

The increasing volume of online discussions requires advanced automatic content moderation to maintain responsible discourse. While hate speech detection on social media is well-studied, research on German-language newspaper forums remains limited. Existing studies often neglect platform-specific context, such as user history and article themes. This paper addresses this gap by developing and evaluating binary classification models for automatic content moderation in German newspaper forums, incorporating contextual information. Using LSTM, CNN, and ChatGPT-3.5 Turbo, and leveraging the One Million Posts Corpus from the Austrian newspaper Der Standard, we assess the impact of context-aware models. Results show that CNN and LSTM models benefit from contextual information and perform competitively with state-of-the-art approaches. In contrast, ChatGPT's zero-shot classification does not improve with added context and underperforms.

内容审核德语NLP上下文建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。