arXiv:2505.18172cs.CRcs.LG2025-05被引 2

提出主动检测框架,防范生成式AI被恶意利用

GenAI Security: Outsmarting the Bots with a Proactive Testing Framework

  • 设计主动测试框架,提前发现生成式AI漏洞
  • 在SPML数据集上验证框架有效,抵御提示注入攻击
  • 适合关注AI安全的开发者与研究者

生成式AI(GenAI)模型日益复杂并广泛集成于各类应用中,带来传统方法难以应对的新安全挑战。本文探讨了针对恶意滥用生成式AI系统风险,亟需采取主动防御措施的必要性。提出一个涵盖关键方法、工具与策略的框架,旨在应对高级对抗性攻击,强调保障生成式AI创新免受潜在法律与安全责任影响的重要性。通过在SPML Chatbot Prompt Injection Dataset上进行实证测试,证明该框架的有效性。本研究凸显了从被动响应转向主动防护的安全范式转变,对生成式AI技术的安全可靠部署至关重要。

原文摘要 · Abstract (English)

The increasing sophistication and integration of Generative AI (GenAI) models into diverse applications introduce new security challenges that traditional methods struggle to address. This research explores the critical need for proactive security measures to mitigate the risks associated with malicious exploitation of GenAI systems. We present a framework encompassing key approaches, tools, and strategies designed to outmaneuver even advanced adversarial attacks, emphasizing the importance of securing GenAI innovation against potential liabilities. We also empirically prove the effectiveness of the said framework by testing it against the SPML Chatbot Prompt Injection Dataset. This work highlights the shift from reactive to proactive security practices essential for the safe and responsible deployment of GenAI technologies

AI安全生成式AI主动检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。