arXiv:2604.00323cs.CLcs.CY2026-04被引 1

LLM如何贯穿网络滥用检测全流程,提升安全系统智能化水平

Large Language Models in the Abuse Detection Pipeline

  • 将大模型用于检测流程四阶段:标注、识别、申诉审核与治理审计
  • 支持政策理解与解释生成,应对复杂多变的滥用行为模式
  • 适合平台安全团队与研究者参考,尤其关注可解释性与鲁棒性

在线滥用行为日益复杂,涵盖毒性语言、骚扰、操纵和欺诈等。传统依赖静态分类器与人工标注的机器学习方法难以跟上威胁演变与政策细微变化。大语言模型(LLMs)具备上下文推理、政策解读、解释生成与跨模态理解能力,可支持现代安全系统的多个环节。本文从生命周期视角分析LLM在滥用检测生命周期(ADL)中的应用,涵盖四个阶段:(I) 标注与特征生成,(II) 检测,(III) 审核与申诉,(IV) 审计与治理。针对每阶段综述前沿研究与产业实践,讨论生产部署架构设计,并评估基于LLM方法的优势与局限。最后指出延迟、成本效率、确定性、对抗鲁棒性与公平性等关键挑战,提出未来需推动大模型成为可信赖、可问责的大规模滥用检测与治理系统核心组件的研究方向。

原文摘要 · Abstract (English)

Online abuse has grown increasingly complex, spanning toxic language, harassment, manipulation, and fraudulent behavior. Traditional machine-learning approaches dependent on static classifiers and labor-intensive labeling struggle to keep pace with evolving threat patterns and nuanced policy requirements. Large Language Models introduce new capabilities for contextual reasoning, policy interpretation, explanation generation, and cross-modal understanding, enabling them to support multiple stages of modern safety systems. This survey provides a lifecycle-oriented analysis of how LLMs are being integrated into the Abuse Detection Lifecycle (ADL), which we define across four stages: (I) Label \& Feature Generation, (II) Detection, (III) Review \& Appeals, and (IV) Auditing \& Governance. For each stage, we synthesize emerging research and industry practices, highlight architectural considerations for production deployment, and examine the strengths and limitations of LLM-driven approaches. We conclude by outlining key challenges including latency, cost-efficiency, determinism, adversarial robustness, and fairness and discuss future research directions needed to operationalize LLMs as reliable, accountable components of large-scale abuse-detection and governance systems.

大模型内容安全滥用检测生命周期

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。