用任务先验重新定义模型评估,覆盖所有可能下游任务。
Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
- 构建任务概率空间,用先验分布替代固定基准集
- 可计算模型在全部任务上的平均表现与方差
- 适合自监督学习研究者快速评估模型泛化能力
人工智能研究的终极目标是让系统能解决任何可能的任务。然而当前评估方法仍依赖于人工挑选的一组固定下游基准。我们提出任务先验(Task Priors)框架,通过定义任务分布来构建下游任务的概率空间,从而实现对模型在所有可能任务上表现的量化评估。该框架首次回答了关键问题:(i)模型在所有可能任务中的加权平均性能如何?(ii)其性能在任务空间中的方差是多少?这为自监督学习(SSL)提供了新的评估标准,有望加速研究进展。
原文摘要 · Abstract (English)
The grand goal of AI research, and particularly Self Supervised Learning (SSL), is to produce systems that can successfully solve any possible task. In contrast, current evaluation methods available to AI researchers typically rely on a fixed collection of hand-picked downstream benchmarks. Hence, a large amount of effort is put into designing and searching for large collection of evaluation tasks that can serve as a proxy of our grand goal. We argue that such a rigid evaluation protocol creates a silent bottleneck in AI research. To remedy that, we define a probabilistic space of downstream tasks obtained by adopting a distribution of tasks and by defining Task Priors. Under this view, one can evaluate a model's performance over the set of all possible downstream tasks. Our framework is the first to provide answers to key questions such as (i) what is the average performance of my model over all possible downstream tasks weighted by the probability to encounter each task? or (ii) what is the variance of my model's performance across all downstream tasks under the defined Task Priors? Beyond establishing a new standard for evaluation, we believe that Task Priors will accelerate the pace of research in SSL - where downstream task evaluation is the sole qualitative signal that researchers have access to.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。