arXiv:2608.27462cs.CLcs.AI2026-08

针对隐性仇恨言论设计自适应检测框架,按难易程度分配不同计算资源。

Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech

论文配图:Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech
图 1 · 摘自论文原文
  • 按语言细微差别分三类:浅层、定向、上下文依赖,分别用轻量、增强、智能生成应对。
  • 在四个数据集上超越当前最佳模型,复杂样本检测准确率提升显著。
  • 适合需要高效精准识别隐性攻击的平台内容审核团队使用。

与明显辱骂不同,隐性仇恨言论通过隐喻和上下文线索隐藏恶意,使在线内容审查难以识别。现有基于PLM或LLM的方法通常对所有样本采用统一推理流程,忽略语言细微差别,导致简单样本产生冗余计算。我们发现网络仇恨言论并非单一形态,而是呈现多样化特征。因此定义三类细粒度类型:浅层、定向、上下文依赖。提出细粒度自适应隐性仇恨言论检测(FAID)框架:先分类,再适配。对表面意图明显的浅层样本,采用轻量提示调优快速判断;对隐含目标的定向评论,设计知识增强机制迭代揭示隐藏目标;对缺乏背景信息的上下文依赖型评论,采用代理式框架自动生成提示,演化上下文、推断缺失信息并识别模糊恶意意图。该架构将计算资源集中于复杂样本,避免对简单样本过度推理。在四个基准数据集上的实验表明,FAID显著优于现有最先进方法。

原文摘要 · Abstract (English)

Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples. This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases. We observe that online hate speech is not monolithic but manifests in varied forms. We therefore define three fine-grained categories: Shallow, Targeted, and Context-Dependent. Accordingly, we propose Fine-grained Adaptive Implicit Hate speech Detection (FAID), a novel framework that first performs fine-grained classification and then adapts to specific categories. Specifically, for Shallow samples with surface-identifiable intents, the framework adopts lightweight prompt-tuning for rapid classification; for Targeted comments that bind malicious intent to concealed targets, we design knowledge augmentation to iteratively refine the model and reveal hidden targets; for Context-Dependent comments lacking background information, we utilize an agentic framework that automatically generates prompts to evolve context, infer missing background information and identify ambiguous malicious intents. This adaptive architecture focuses computational resources on complex implicit samples while avoiding redundant reasoning for shallow samples. Experiments on four benchmark datasets demonstrate that FAID significantly outperforms SOTA baselines.

隐性仇恨自适应检测大模型应用内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。