通过判别分解识别优质查询集,提升黑箱模型分类精度
Black-box model classification under the discriminative factorization
- 提出判别分解方法,区分高质量与低质量查询集
- 查询预算增加时,随机分类概率呈指数级下降
- 适合需要评估黑箱模型性能的研究者使用
现代生成系统常以API形式提供(即“黑箱”场景),用户在推理时无法获知模型内部属性。尽管已有研究证明,基于一组查询下模型嵌入响应关系的低维表征可用于推断模型级别属性,但此类表征质量高度依赖查询集选择。本文引入判别分解框架,用于在黑箱模型分类任务中区分高质量与低质量查询集。在此框架下,随机分类概率随查询预算增加呈指数衰减。在三个审计任务中,估计的分解参数能准确预测实际性能衰减速率。最终表明,利用估计的判别场选择的查询集可复现最优查询集的排序效果。
原文摘要 · Abstract (English)
Access to modern generative systems is often restricted to querying an API (the ``black-box" setting) and many properties of the system are unknown to the user at inference time. While recent work has shown that low-dimensional representations of models based on the relationship between their embedded responses to a set of queries are useful for inferring model-level properties, the quality of these representations is highly sensitive to the query set. We introduce the \emph{discriminative factorization} to distinguish between high- and low-quality query sets in the context of black-box model-level classification. Under this framework, the probability of chance-level classification decays exponentially in the query budget. On three auditing tasks, estimated factorization parameters predict the empirical performance decay rate. We conclude by showing that query sets selected using the estimated discriminative field reproduce the empirical ordering of oracle query sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。