arXiv:2505.04784cs.CRcs.AI2025-05被引 1

提出多维度风险评估框架,量化大模型聊天机器人对三方的潜在威胁。

A Proposal for Evaluating the Operational Risk for ChatBots based on Large Language Models

  • 基于技术复杂度与上下文因素,评估聊天机器人漏洞风险。
  • 通过增强版Garak工具,覆盖误导、代码幻觉等五类威胁向量。
  • 适用于需保障安全可靠的AI对话系统的研发与运维团队。

生成式AI与大语言模型的兴起使聊天机器人具备类人交互能力,但其引入的操作风险远超传统网络安全范畴。本文提出一种新型可量化的风险评估指标,同时评估服务提供方、终端用户及第三方三类主体面临的风险。该方法融合诱发错误行为的技术复杂度(从非诱导故障到高级提示注入攻击)以及目标行业、用户年龄范围、漏洞严重性等上下文因素。为验证该指标,我们采用开源的Garak框架进行测试,并进一步扩展其功能以捕获包括误导信息、代码幻觉、社交工程和恶意代码生成在内的多种威胁向量。在使用检索增强生成(RAG)的聊天机器人场景中,聚合风险评分有效指导了短期缓解措施与长期模型设计优化。结果表明,多维度风险评估对实现安全可靠的AI驱动对话系统至关重要。

原文摘要 · Abstract (English)

The emergence of Generative AI (Gen AI) and Large Language Models (LLMs) has enabled more advanced chatbots capable of human-like interactions. However, these conversational agents introduce a broader set of operational risks that extend beyond traditional cybersecurity considerations. In this work, we propose a novel, instrumented risk-assessment metric that simultaneously evaluates potential threats to three key stakeholders: the service-providing organization, end users, and third parties. Our approach incorporates the technical complexity required to induce erroneous behaviors in the chatbot--ranging from non-induced failures to advanced prompt-injection attacks--as well as contextual factors such as the target industry, user age range, and vulnerability severity. To validate our metric, we leverage Garak, an open-source framework for LLM vulnerability testing. We further enhance Garak to capture a variety of threat vectors (e.g., misinformation, code hallucinations, social engineering, and malicious code generation). Our methodology is demonstrated in a scenario involving chatbots that employ retrieval-augmented generation (RAG), showing how the aggregated risk scores guide both short-term mitigation and longer-term improvements in model design and deployment. The results underscore the importance of multi-dimensional risk assessments in operationalizing secure, reliable AI-driven conversational systems.

大模型安全风险评估聊天机器人LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。