提出新方法评估大模型在输入扰动下的稳定性,无需修改模型参数。
SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- 基于图流形间的距离映射畸变度量,实现样本级鲁棒性分析。
- 在攻击效率和鲁棒训练上均有显著提升,适用于大模型与小模型。
- 无需复杂扰动设计,可作为通用工具提升NLP系统可靠性。
近年来,基于预训练Transformer的语言模型在众多自然语言处理任务中取得了领先性能。然而,随着模型规模和部署范围的扩大,其对输入扰动的鲁棒性成为紧迫问题。现有鲁棒性方法在小参数模型与大规模语言模型(LLMs)之间存在分歧,且通常依赖耗时的、样本特定的对抗性设计。本文提出统一的局部(样本级)鲁棒性框架SALMAN,无需修改内部参数或使用复杂扰动启发式策略即可评估模型稳定性。核心是新颖的距离映射畸变(DMD)度量,通过近线性复杂度比较输入-输出距离映射,对每个样本的敏感性进行排序。实验表明,该框架在攻击效率和鲁棒训练方面均有显著提升,可作为模型无关的实用工具,推动Transformer-based NLP系统的可靠性发展。
原文摘要 · Abstract (English)
Recent strides in pretrained transformer-based language models have propelled state-of-the-art performance in numerous NLP tasks. Yet, as these models grow in size and deployment, their robustness under input perturbations becomes an increasingly urgent question. Existing robustness methods often diverge between small-parameter and large-scale models (LLMs), and they typically rely on labor-intensive, sample-specific adversarial designs. In this paper, we propose a unified, local (sample-level) robustness framework (SALMAN) that evaluates model stability without modifying internal parameters or resorting to complex perturbation heuristics. Central to our approach is a novel Distance Mapping Distortion (DMD) measure, which ranks each sample's susceptibility by comparing input-to-output distance mappings in a near-linear complexity manner. By demonstrating significant gains in attack efficiency and robust training, we position our framework as a practical, model-agnostic tool for advancing the reliability of transformer-based NLP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。