arXiv:2605.10563cs.CLcs.AI2026-05

构建细粒度威胁检测基准,区分显性、隐性威胁与非威胁

ThreatCore: A Benchmark for Explicit and Implicit Threat Detection

论文配图:ThreatCore: A Benchmark for Explicit and Implicit Threat Detection
图 1 · 摘自论文原文
  • 统一定义威胁,重标注多源数据并生成合成样本提升覆盖
  • 隐性威胁检测准确率显著低于显性威胁,现有模型仍难应对
  • 引入语义角色标注可提升模型表现,适合安全与内容审核研究者

自然语言处理中的威胁检测缺乏一致定义和标准基准,常与毒性、仇恨言论等概念混淆。本文提出ThreatCore,一个公开可用的细粒度威胁检测基准数据集,明确区分显性威胁、隐性威胁和非威胁。数据集通过聚合多个公开资源并基于统一操作定义重新标注,揭示了现有标签存在显著不一致。为提升隐性威胁等少数类覆盖,我们采用人工验证的合成样本进行增强,确保各数据源一致性。在ThreatCore上评估Perspective API、零样本分类器及近期语言模型,结果表明隐性威胁检测难度远高于显性威胁。同时,引入语义角色标注作为中间表示,可使有害意图结构更清晰,性能提升明显。ThreatCore为细粒度威胁检测提供了更一致的基准,并凸显当前模型在识别间接危害意图方面的挑战。

原文摘要 · Abstract (English)

Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomena such as toxicity, hate speech, or offensive language. In this work, we introduce ThreatCore, a public available benchmark dataset for fine-grained threat detection that distinguishes between explicit threats, implicit threats, and non-threats. The dataset is constructed by aggregating multiple publicly available resources and systematically re-annotating them under a unified operational definition of threat, revealing substantial inconsistencies across existing labels. To improve the coverage of underrepresented cases, particularly implicit threats, we further augment the dataset with synthetic examples, which are manually validated using the same annotation protocol adopted for the re-annotation of the public datasets, ensuring consistency across all data sources. We evaluate Perspective API, zero-shot classifiers, and recent language models on ThreatCore, showing that implicit threats remain substantially harder to detect than explicit ones. Our results also indicate that incorporating Semantic Role Labeling as an intermediate representation can improve performance by making the structure of harmful intent more explicit. Overall, ThreatCore provides a more consistent benchmark for studying fine-grained threat detection and highlights the challenges that current models still face in identifying indirect expressions of harmful intent.

威胁检测细粒度分析语义角色标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。