arXiv:2510.05440stat.MLcs.CR2025-10被引 1

通过两个对手证明者,用极少查询实现高精度模型选择。

Refereed Learning

  • 设计双证明者协议,仅需一次真值查询即可评估黑箱模型。
  • 在高精度场景下,通信量为(1+1/ε²)·poly(d),误差小于(1+ε)。
  • 适用于需要极低标注成本的模型筛选任务,如自动化机器学习。

我们首次研究在存在两个竞争证明者(仅一个诚实)的设置下的学习任务,重点评估黑箱模型的真实性能。该框架称为裁判学习(refereed learning)。我们提出通用定义并设计协议,在仅使用一次真值函数查询、通信量为(1+1/ε²)·poly(d)比特的条件下,输出模型的损失与最优模型相差不超过(1+ε)倍。而单个证明者需几乎全量访问真值点才能达到类似精度。我们还给出了下界,证明协议在证明者复杂度、样本数和查询需求上均接近最优。

原文摘要 · Abstract (English)

We initiate an investigation of learning tasks in a setting where the learner is given access to two competing provers, only one of which is honest. Specifically, we consider the power of such learners in assessing purported properties of opaque models. Following prior work in complexity theory that considers the power of competing provers in various settings, we call this setting refereed learning. After formulating a general definition of refereed learning tasks, we show refereed learning protocols that obtain a level of accuracy that far exceeds what is obtainable at comparable cost without provers, or even with a single prover. We concentrate on the task of choosing the better one out of two black-box models, with respect to some ground truth. While we consider a range of parameters, perhaps our most notable result is in the high-precision range: For all $\varepsilon>0$ and ambient dimension $d$, our learner makes only one query to the ground truth function, communicates only $(1+\frac{1}{\varepsilon^2})\cdot\text{poly}(d)$ bits with the provers, and outputs a model whose loss is within a multiplicative factor of $(1+\varepsilon)$ of the best model's loss. Obtaining comparable loss with a single prover would require the learner to access the ground truth at almost all of the points in the domain. We also present lower bounds that demonstrate the optimality of our protocols in a number of respects, including prover complexity, number of samples, and need for query access.

模型评估对抗学习高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。