arXiv:2506.05203cs.LOcs.LG2025-06被引 4

验证机器学习副本是否保持可信性,提出四类可信度判定标准。

Trustworthiness Preservation by Copies of Machine Learning Systems

  • 通过形式化计算模型分析副本与原系统的行为差异
  • 定义了四类可信度:正当、等同、弱、几乎可信,可逐层验证
  • 适用于需保证一致性与责任的AI系统复制场景

机器学习系统开发中常在不同数据集上训练相同模型,或用相同数据集训练多个模型,这些系统被视为原系统的副本。在负责任AI背景下,一个关键但少被研究的问题是:如何验证副本是否保持了原系统的可信性?本文提出一种形式化计算框架,用于建模和验证数据上的概率复杂查询,并定义了四种可信度概念:正当可信、等同可信、弱可信、几乎可信。这些概念可通过对比副本与原系统的行为进行检验。论文进一步研究了各类可信度之间的关系及其在逻辑运算下的组合性质,旨在为已知行为的原始系统提供一套可计算的工具,以验证其复杂副本的可信性。

原文摘要 · Abstract (English)

A common practice of ML systems development concerns the training of the same model under different data sets, and the use of the same (training and test) sets for different learning models. The first case is a desirable practice for identifying high quality and unbiased training conditions. The latter case coincides with the search for optimal models under a common dataset for training. These differently obtained systems have been considered akin to copies. In the quest for responsible AI, a legitimate but hardly investigated question is how to verify that trustworthiness is preserved by copies. In this paper we introduce a calculus to model and verify probabilistic complex queries over data and define four distinct notions: Justifiably, Equally, Weakly and Almost Trustworthy which can be checked analysing the (partial) behaviour of the copy with respect to its original. We provide a study of the relations between these notions of trustworthiness, and how they compose with each other and under logical operations. The aim is to offer a computational tool to check the trustworthiness of possibly complex systems copied from an original whose behavour is known.

可信性机器学习形式验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。