无需标注数据,用模型自评实现任务定制的可信排名。
CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks
- 多模型互评:轮流扮演教师、学生和裁判,生成无污染基准。
- 在真实标签下相关性达0.86,十三模型排名斯皮尔曼相关0.95。
- 自动剔除无效裁判,适合个性化任务选型与模型迭代。
为特定应用选择预训练语言模型或评估微调效果是一项高价值决策,但现有公共基准存在缺陷:通用基准未必反映具体子领域或子任务,且当测试题泄露至预训练数据时,得分可能源于记忆而非真正求解。本文提出 CoEval,一个开放框架,通过集成自评估实现可信、任务定制的信号:给定任务或领域描述,一组模型轮流担任教师、学生和裁判,生成全新、无污染的基准,回答并互相评分,全程无需人工标注。每个模型作为学生参与答题,其响应用于加权问题(按区分能力)和裁判(按群体共识)。当存在真实答案时,CoEval 在 {ho}=0.86 时恢复正确排序;对十三个模型的排名,斯皮尔曼相关达 0.95。可靠性来自评委组合,而非数量:无标签加权机制可剔除失效裁判、降低饱和问题权重,避免偏差。生成题目与五个公开基准无文字重复;评委组消除冗余偏见,杜绝同家族模型自偏爱;排名具有领域特异性:在四个新领域中,不同模型各居榜首,说明通用排行榜误导多数从业者。该流程可每次模型发布后重跑,为团队提供无污染的应用专属排行榜。
原文摘要 · Abstract (English)
Selecting a pretrained language model, or evaluating a fine-tuned one, for a specific application is a high-value decision, yet the public benchmarks used to make it are poorly suited: a generic benchmark need not reflect a particular sub-domain or sub-task, and its scores are suspect when its items have leaked into pretraining and are recalled rather than solved. We present CoEval, an open framework that supplies a trustworthy, task-specific signal through ensemble self-evaluation: from a task or domain description, a pool of models rotates through all three roles, teacher, student, and judge, to generate a fresh, contamination-free benchmark, answer it, and score one another, with no human labels or raters. Because every model also answers as a student, the responses are the data that weight each question by its discriminative power and each judge by its consensus with the panel. Where ground truth exists, CoEval recovers the true ranking and tracks objective correctness at \r{ho}=0.86, and the weighting recovers the gold ranking of thirteen models at Spearman 0.95. Reliability comes from panel composition, not size: this label-free weighting zeroes out broken judges and down-weights saturated questions, so neither distorts the ranking. Generated items show zero verbatim overlap with five public benchmarks, the panel cancels verbosity bias and precludes same-family self-preference, and rankings are domain-specific: three different models top four de-novo domains, so a generic leaderboard misdirects most practitioners. The same pipeline reruns on each model release, giving any team a contamination-free leaderboard for its application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。