arXiv:2512.12583cs.CRcs.AI2025-12被引 1

用分类器检测大模型应用中的恶意提示注入,提升系统安全性。

Detecting Prompt Injection Attacks Against Application Using Classifiers

  • 基于真实攻击数据构建并扩充提示注入数据集
  • 多种模型在检测准确率上优于基线方法
  • 适合关注LLM安全的开发者与安全研究人员

提示注入攻击可能危及从基础设施到大型网络应用的关键系统安全与稳定性。本文基于HackAPrompt Playground提交数据集构建并扩充了提示注入数据集,训练了LSTM、前馈神经网络、随机森林和朴素贝叶斯等多种分类器,用于检测集成大模型的Web应用中的恶意提示。所提方法显著提升了提示注入的检测与缓解能力,有助于保护目标应用与系统。

原文摘要 · Abstract (English)

Prompt injection attacks can compromise the security and stability of critical systems, from infrastructure to large web applications. This work curates and augments a prompt injection dataset based on the HackAPrompt Playground Submissions corpus and trains several classifiers, including LSTM, feed forward neural networks, Random Forest, and Naive Bayes, to detect malicious prompts in LLM integrated web applications. The proposed approach improves prompt injection detection and mitigation, helping protect targeted applications and systems.

提示注入LLM安全分类器威胁检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。