探究模型表示相似性的成因及其安全风险。
Causes and Consequences of Representational Similarity in Machine Learning Models
- 通过数据集和任务重叠度实验,分析模型相似性来源。
- 重叠度越高,模型表示越相似,大模型效应更明显。
- 相似性提升使模型更易遭受迁移攻击,需警惕安全风险。
众多研究发现机器学习模型在不同模态间存在表示上的相似性。尽管已有大量工作探索模型对齐的性质与度量方法,但对其成因的研究仍很有限。本文系统考察了数据集重叠和任务重叠两个因素对下游模型表示相似性的影响。通过在不同规模和模态(从小型分类器到大型语言模型)上的实验,我们发现两者均会提高表示相似性,且联合效应最强。进一步分析表明,更高的表示相似性会增加模型对可迁移对抗攻击和越狱攻击的脆弱性,揭示了其潜在的安全后果。
原文摘要 · Abstract (English)
Numerous works have noted similarities in how machine learning models represent the world, even across modalities. Although much effort has been devoted to uncovering properties and metrics on which these models align, surprisingly little work has explored causes of this similarity. To advance this line of inquiry, this work explores how two factors - dataset overlap and task overlap - influence downstream model similarity. We evaluate the effects of both factors through experiments across model sizes and modalities, from small classifiers to large language models. We find that both task and dataset overlap cause higher representational similarity and that combining them provides the strongest effect. Finally, we consider downstream consequences of representational similarity, demonstrating how greater similarity increases vulnerability to transferable adversarial and jailbreak attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。