arXiv:2512.17367cs.LG2025-12

提出新框架与检测器,提升网络有害内容识别在对抗攻击下的鲁棒性。

Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach

  • 基于大模型生成对抗样本并聚合,增强检测器对多样攻击的泛化能力。
  • 在三个数据集上,对抗条件下准确率显著提升,且对多种攻击类型保持稳定。
  • 适合安全团队、平台方及关注对抗鲁棒性的研究人员使用。

社交媒体充斥着仇恨言论、虚假信息和极端主义言论等有害内容。尽管机器学习模型被广泛用于检测,但极易受对抗攻击影响——恶意用户通过细微修改文本即可逃避检测。提升对抗鲁棒性至关重要,需同时具备强泛化性与高准确率,但二者难以兼顾。本研究遵循计算设计科学范式,首先提出一种新框架(LLM-SGA),通过识别文本对抗攻击的关键不变性,利用大语言模型生成并聚合样本,确保检测器具有强泛化能力。其次,构建新型检测器(ARHOCD),包含三项创新:(1)多基检测器集成,发挥互补优势;(2)基于可预测性与检测器能力动态调整权重,初始值基于领域知识,通过贝叶斯推断更新;(3)迭代优化基检测器与权重分配器的对抗训练策略。针对现有研究局限,我们在涵盖仇恨言论、谣言与极端内容的三个数据集上进行实证评估。结果表明,ARHOCD在对抗环境下兼具强泛化性和高准确率。

原文摘要 · Abstract (English)

Social media platforms are plagued by harmful content such as hate speech, misinformation, and extremist rhetoric. Machine learning (ML) models are widely adopted to detect such content; however, they remain highly vulnerable to adversarial attacks, wherein malicious users subtly modify text to evade detection. Enhancing adversarial robustness is therefore essential, requiring detectors that can defend against diverse attacks (generalizability) while maintaining high overall accuracy. However, simultaneously achieving both optimal generalizability and accuracy is challenging. Following the computational design science paradigm, this study takes a sequential approach that first proposes a novel framework (Large Language Model-based Sample Generation and Aggregation, LLM-SGA) by identifying the key invariances of textual adversarial attacks and leveraging them to ensure that a detector instantiated within the framework has strong generalizability. Second, we instantiate our detector (Adversarially Robust Harmful Online Content Detector, ARHOCD) with three novel design components to improve detection accuracy: (1) an ensemble of multiple base detectors that exploits their complementary strengths; (2) a novel weight assignment method that dynamically adjusts weights based on each sample's predictability and each base detector's capability, with weights initialized using domain knowledge and updated via Bayesian inference; and (3) a novel adversarial training strategy that iteratively optimizes both the base detectors and the weight assignor. We addressed several limitations of existing adversarial robustness enhancement research and empirically evaluated ARHOCD across three datasets spanning hate speech, rumor, and extremist content. Results show that ARHOCD offers strong generalizability and improves detection accuracy under adversarial conditions.

对抗鲁棒内容检测大模型安全防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。