无需标签数据,用大模型预测错误来可靠比较跨领域模型性能。
Can We Reliably Rank Model Performance across Domains without Labeled Data?
- 用大语言模型预测错误,替代传统相似性或漂移方法。
- 在15个领域上,相关性比基线高30%以上,结果更稳定。
- 适合做跨领域模型评估,尤其当性能差异明显时
在无标签条件下估计模型性能对理解NLP模型泛化能力至关重要。本文通过两阶段评估框架,使用4个基础分类器和多个大语言模型作为错误预测器,在GeoOLID与Amazon Reviews数据集(共15个领域)上进行实验。结果表明,基于大语言模型的错误预测器产生的性能排名相关性显著优于基于数据漂移或零样本的基线方法。分析发现:当不同领域间性能差异较大,且错误预测模型与基础模型的真实失败模式一致时,排名可靠性更高。研究厘清了性能估计方法的可信边界,为跨领域模型评估提供实用指导。
原文摘要 · Abstract (English)
Estimating model performance without labels is an important goal for understanding how NLP models generalize. While prior work has proposed measures based on dataset similarity or predicted correctness, it remains unclear when these estimates produce reliable performance rankings across domains. In this paper, we analyze the factors that affect ranking reliability using a two-step evaluation setup with four base classifiers and several large language models as error predictors. Experiments on the GeoOLID and Amazon Reviews datasets, spanning 15 domains, show that large language model-based error predictors produce stronger and more consistent rank correlations with true accuracy than drift-based or zero-shot baselines. Our analysis reveals two key findings: ranking is more reliable when performance differences across domains are larger, and when the error model's predictions align with the base model's true failure patterns. These results clarify when performance estimation methods can be trusted and provide guidance for their use in cross-domain model evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。