arXiv:2502.07036cs.CRcs.AI2025-02被引 11

提出LLM一致性评估框架,发现主流模型在安全问答中常自相矛盾。

Automated Consistency Analysis of LLMs

  • 定义响应一致性并构建评估框架,支持自检与多模型对比验证
  • 测试GPT4oMini、GPT3.5等模型在安全问答任务中一致性不足
  • 适用于安全领域可信AI评估,帮助筛选可靠LLM应用

生成式人工智能(Gen AI)与大型语言模型(LLMs)已在工业、学术及政府领域广泛应用。网络安全是其重要应用领域之一。然而,可信的生成式AI与LLMs在关键领域的部署仍面临诸多挑战。其中核心问题在于:大模型响应的一致性如何?本文首次对LLM响应一致性进行形式化定义,并构建了评估框架。提出两种验证方法:自我验证与跨模型验证。在包含多个网络安全问题的信息性与情境性任务基准上,对GPT4oMini、GPT3.5、Gemini、Cohere和Llama3等模型进行了广泛实验。结果表明,尽管这些模型已被用于或正被考虑用于多种网络安全任务,其响应往往缺乏一致性,因此在网络安全场景中不可靠、不值得信赖。

原文摘要 · Abstract (English)

Generative AI (Gen AI) with large language models (LLMs) are being widely adopted across the industry, academia and government. Cybersecurity is one of the key sectors where LLMs can be and/or are already being used. There are a number of problems that inhibit the adoption of trustworthy Gen AI and LLMs in cybersecurity and such other critical areas. One of the key challenge to the trustworthiness and reliability of LLMs is: how consistent an LLM is in its responses? In this paper, we have analyzed and developed a formal definition of consistency of responses of LLMs. We have formally defined what is consistency of responses and then develop a framework for consistency evaluation. The paper proposes two approaches to validate consistency: self-validation, and validation across multiple LLMs. We have carried out extensive experiments for several LLMs such as GPT4oMini, GPT3.5, Gemini, Cohere, and Llama3, on a security benchmark consisting of several cybersecurity questions: informational and situational. Our experiments corroborate the fact that even though these LLMs are being considered and/or already being used for several cybersecurity tasks today, they are often inconsistent in their responses, and thus are untrustworthy and unreliable for cybersecurity.

LLM评估一致性网络安全可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。