探究大模型在对抗攻击与分布外输入下的鲁棒性关系,发现二者难以互相提升。
On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models
- 跨场景测试对抗与分布外鲁棒性方法的迁移效果。
- 小模型相关性中性,大模型负相关,Mixtral呈正相关。
- 提示需针对模型类型定制混合鲁棒性策略。
大型语言模型(LLMs)在各类应用中的依赖日益增加,亟需理解其对对抗扰动和分布外(OOD)输入的鲁棒性。本文研究了LLMs中对抗鲁棒性与OOD鲁棒性之间的关联,填补了鲁棒性评估的关键空白。通过将原本用于提升一种鲁棒性的方法应用于两种场景,我们分析其在对抗与分布外基准数据集上的表现。模型输入为文本样本,输出预测以自然语言推理任务中的准确率、精确率、召回率和F1分数进行评估。结果揭示了二者之间复杂的相互作用:两种鲁棒性类型间转移能力有限。通过针对性消融实验,我们发现其相关性随模型规模和架构变化而呈现模型特异性趋势:较小模型如LLaMA2-7b表现为中性相关,较大模型如LLaMA2-13b显示负相关,而Mixtral则表现出正相关,可能源于领域特定对齐。这些结果强调了整合对抗与分布外策略的混合鲁棒性框架的重要性,需根据具体模型与领域定制。未来研究需在更大模型与多样化架构上进一步评估此类交互,以推动更可靠、泛化性更强的LLMs发展。
原文摘要 · Abstract (English)
The increasing reliance on large language models (LLMs) for diverse applications necessitates a thorough understanding of their robustness to adversarial perturbations and out-of-distribution (OOD) inputs. In this study, we investigate the correlation between adversarial robustness and OOD robustness in LLMs, addressing a critical gap in robustness evaluation. By applying methods originally designed to improve one robustness type across both contexts, we analyze their performance on adversarial and out-of-distribution benchmark datasets. The input of the model consists of text samples, with the output prediction evaluated in terms of accuracy, precision, recall, and F1 scores in various natural language inference tasks. Our findings highlight nuanced interactions between adversarial robustness and OOD robustness, with results indicating limited transferability between the two robustness types. Through targeted ablations, we evaluate how these correlations evolve with different model sizes and architectures, uncovering model-specific trends: smaller models like LLaMA2-7b exhibit neutral correlations, larger models like LLaMA2-13b show negative correlations, and Mixtral demonstrates positive correlations, potentially due to domain-specific alignment. These results underscore the importance of hybrid robustness frameworks that integrate adversarial and OOD strategies tailored to specific models and domains. Further research is needed to evaluate these interactions across larger models and varied architectures, offering a pathway to more reliable and generalizable LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。