现有漏洞评分系统难评估大模型对抗攻击,因评分变化小且规则僵化。
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
- 用3个大模型平均打分,对比56种对抗攻击的漏洞分
- 多数攻击得分差异极小,说明评分体系无法区分攻击强弱
- 适合关注大模型安全评估、需改进评分标准的研究者
本研究检验了通用漏洞评分系统(CVSS)等传统指标在评估大语言模型(LLMs)对抗攻击(AAs)时的有效性。通过量化分析56种来自不同论文和在线数据库的对抗攻击,采用三个大模型分别评分后取平均,计算漏洞评分的变异系数。结果显示,现有评分体系对不同攻击的评分差异极小,表明其内部因素(尤其是具有固定取值的上下文相关因子)难以有效区分攻击威胁程度。这验证了当前基于固定规则的漏洞评分系统在评估大模型对抗攻击时存在局限,亟需开发更灵活、通用的新评估框架。研究为提升大模型安全评估体系提供了实证依据。
原文摘要 · Abstract (English)
This research investigates the effectiveness of established vulnerability metrics, such as the Common Vulnerability Scoring System (CVSS), in evaluating attacks against Large Language Models (LLMs), with a focus on Adversarial Attacks (AAs). The study explores the influence of both general and specific metric factors in determining vulnerability scores, providing new perspectives on potential enhancements to these metrics. This study adopts a quantitative approach, calculating and comparing the coefficient of variation of vulnerability scores across 56 adversarial attacks on LLMs. The attacks, sourced from various research papers, and obtained through online databases, were evaluated using multiple vulnerability metrics. Scores were determined by averaging the values assessed by three distinct LLMs. The results indicate that existing scoring-systems yield vulnerability scores with minimal variation across different attacks, suggesting that many of the metric factors are inadequate for assessing adversarial attacks on LLMs. This is particularly true for context-specific factors or those with predefined value sets, such as those in CVSS. These findings support the hypothesis that current vulnerability metrics, especially those with rigid values, are limited in evaluating AAs on LLMs, highlighting the need for the development of more flexible, generalized metrics tailored to such attacks. This research offers a fresh analysis of the effectiveness and applicability of established vulnerability metrics, particularly in the context of Adversarial Attacks on Large Language Models, both of which have gained significant attention in recent years. Through extensive testing and calculations, the study underscores the limitations of these metrics and opens up new avenues for improving and refining vulnerability assessment frameworks specifically tailored for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。