arXiv:2502.03937cs.LGcs.NA2025-02

多模型同时出错风险高,尤其当它们共享算法或数据时。

Quantifying Correlations of Machine Learning Models

  • 通过真实数据模拟三种错误相关场景,量化模型间误差关联度。
  • 共享算法、训练数据或基础模型的模型,错误相关性显著增强。
  • 适用于关注模型安全与系统可靠性的研究者和工程师。

机器学习模型被广泛应用于安全关键领域,其错误可能对用户造成伤害。当多个模型并行部署并同时出错时,风险会加剧。本文探讨了三种导致多模型错误相关性的场景,并利用真实数据进行仿真,量化不同模型间的错误相关性。结果表明,当模型共享相似算法、训练数据或基础模型时,累积风险显著升高。总体来看,模型间的相关性普遍存在,且随着对基础模型和公共数据集依赖的增加而进一步强化,凸显了制定有效缓解策略的必要性。

原文摘要 · Abstract (English)

Machine Learning models are being extensively used in safety critical applications where errors from these models could cause harm to the user. Such risks are amplified when multiple machine learning models, which are deployed concurrently, interact and make errors simultaneously. This paper explores three scenarios where error correlations between multiple models arise, resulting in such aggregated risks. Using real-world data, we simulate these scenarios and quantify the correlations in errors of different models. Our findings indicate that aggregated risks are substantial, particularly when models share similar algorithms, training datasets, or foundational models. Overall, we observe that correlations across models are pervasive and likely to intensify with increased reliance on foundational models and widely used public datasets, highlighting the need for effective mitigation strategies to address these challenges.

模型安全误差相关可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。