arXiv:2504.03714cs.LGcs.AI2025-04Conference of the …被引 4

发现大模型脆弱性根源,定位易受攻击的参数与输入维度

Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models

论文配图:Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
图 1 · 摘自论文原文
  • 提出基于信息几何的敏感度度量FI,识别关键脆弱点
  • 1.5B至13B参数模型中,少数高FI项主导模型不稳定性
  • 修复脆弱参数可提升模型合并性能,适用于鲁棒性研究

大型语言模型(LLMs)和视觉-语言模型(VLMs)在众多任务中表现卓越,但仍易受精心设计的扰动影响。本研究旨在揭示其脆弱性的来源,通过识别对扰动敏感的参数及输入维度(像素或标记嵌入)。为此,我们提出一种基于信息几何的稳定性度量—— extbf{FI}(一阶局部影响),用于量化单个参数与输入维度的敏感性。我们在从1.5B到13B参数的多种LLMs与VLMs上进行广泛分析,发现:(I) 少数高FI值的参数或输入维度显著贡献于模型脆性;(II) 在模型合并过程中抑制这些脆弱参数的影响,可提升整体性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) and Vision-Language Models (VLMs) have achieved impressive performance across a wide range of tasks, yet they remain vulnerable to carefully crafted perturbations. In this study, we seek to pinpoint the sources of this fragility by identifying parameters and input dimensions (pixels or token embeddings) that are susceptible to such perturbations. To this end, we propose a stability measure called \textbf{FI}, \textbf{F}irst order local \textbf{I}nfluence, which is rooted in information geometry and quantifies the sensitivity of individual parameter and input dimensions. Our extensive analysis across LLMs and VLMs (from 1.5B to 13B parameters) reveals that: (I) A small subset of parameters or input dimensions with high FI values disproportionately contribute to model brittleness. (II) Mitigating the influence of these vulnerable parameters during model merging leads to improved performance.

大模型安全脆弱性分析模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。