arXiv:2510.02956cs.LGcs.CV2025-10被引 1

用置信度和分散度评估模型泛化能力,无需标签也能排序模型。

Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking

  • 利用预测置信度与类别分散度的互补信号评估模型。
  • 混合指标在多种数据分布下表现优于单一指标,核范数效果最稳定。
  • 适合无标签测试数据的模型选型与部署评估场景。

在缺乏标签测试数据的情况下,评估模型在分布偏移下的泛化能力至关重要。本文提出一个统一且实用的无监督模型评估与排序框架,适用于两种常见部署场景:(1) 在多个无标签测试集上估计固定模型的准确率(数据集中心评估);(2) 在单个无标签测试集上对一组候选模型进行排序(模型中心评估)。我们证明,模型预测中的两个内在属性——置信度(反映预测确定性)和分散度(捕捉预测类别的多样性)——共同提供了强而互补的泛化信号。我们在多种模型架构、数据集及分布偏移类型下系统评估了基于置信度、分散度及混合指标的表现。结果表明,混合指标在两类评估设置中均持续优于单方面指标。尤其地,预测矩阵的核范数在各类任务中表现稳健准确,包括真实世界数据集,并在中等类别不平衡下仍保持可靠性。这些发现为部署场景下的无监督模型评估提供了实用且可推广的基础。

原文摘要 · Abstract (English)

Assessing model generalization under distribution shift is essential for real-world deployment, particularly when labeled test data is unavailable. This paper presents a unified and practical framework for unsupervised model evaluation and ranking in two common deployment settings: (1) estimating the accuracy of a fixed model on multiple unlabeled test sets (dataset-centric evaluation), and (2) ranking a set of candidate models on a single unlabeled test set (model-centric evaluation). We demonstrate that two intrinsic properties of model predictions, namely confidence (which reflects prediction certainty) and dispersity (which captures the diversity of predicted classes), together provide strong and complementary signals for generalization. We systematically benchmark a set of confidence-based, dispersity-based, and hybrid metrics across a wide range of model architectures, datasets, and distribution shift types. Our results show that hybrid metrics consistently outperform single-aspect metrics on both dataset-centric and model-centric evaluation settings. In particular, the nuclear norm of the prediction matrix provides robust and accurate performance across tasks, including real-world datasets, and maintains reliability under moderate class imbalance. These findings offer a practical and generalizable basis for unsupervised model assessment in deployment scenarios.

模型评估无监督泛化能力置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。