统一评测框架,让上下文学习分类实验可比。
StaICC: Standardized Evaluation for Classification Task in In-context Learning
- 设计标准化工具StaICC,固定提示模板与数据处理流程。
- 涵盖10个常用数据集,减少因设置差异导致的结果偏差。
- 支持诊断性评估,适合希望公平比较模型性能的研究者。
分类任务在上下文学习(In-Context Learning, ICL)中被广泛研究,但现有工作基于互不重叠的基准和设置,其性能受提示模板、数据采样方式、指令设计等非关键因素显著影响,导致文献间结果不一致,难以进行公平比较或元分析。为此,本文提出标准化、易用的评测工具StaICC,用于上下文分类任务。针对常规分类任务,提供StaICC-Normal,包含10个常用数据集,并采用固定格式生成提示,以降低实验实现差异。为进一步扩展应用,还构建子基准StaICC-Diag,用于从多个维度诊断ICL表现,提升推理鲁棒性。
原文摘要 · Abstract (English)
Classification tasks are widely investigated in the In-Context Learning (ICL) paradigm. However, current efforts are evaluated on disjoint benchmarks and settings, while their performances are significantly influenced by some trivial variables, such as prompt templates, data sampling, instructions, etc., which leads to significant inconsistencies in the results reported across various literature, preventing fair comparison or meta-analysis across different papers. Therefore, this paper proposes a standardized and easy-to-use evaluation toolkit (StaICC) for in-context classification. Including, for the normal classification task, we provide StaICC-Normal, selecting 10 widely used datasets, and generating prompts with a fixed form, to mitigate the variance among the experiment implementations. To enrich the usage of our benchmark, we also provide a sub-benchmark StaICC-Diag for diagnosing ICL from several aspects, aiming for a more robust inference processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。