用封闭式竞赛评估大模型,防止刷分,提升结果可信度
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
- 赛题密封至评测时开放,提交前冻结,防提前准备
- 统一评测环境运行所有模型,确保公平性与可复现性
- 赛后公开题目和代码,便于社区审计与学习
当前自然语言处理的进展主要通过基准测试和排行榜来体现,但在大模型时代,这些指标越来越容易被误读。得分可能反映的是对特定榜单的迎合、隐藏的评估选择或意外接触到测试内容,而非真正的综合能力。封闭式基准虽能缓解部分问题,但降低了透明度,阻碍了社区从结果中学习。我们提出一种互补的实践:采用奥林匹克风格的评测活动,即在评测前密封题目,提交内容提前冻结,所有模型通过统一标准化的评测流程运行。评分完成后,完整任务集和评估代码将公开,以支持结果的复现与审计。该设计旨在使高性能更难人为制造,同时更易于信任。
原文摘要 · Abstract (English)
Benchmarks and leaderboards are how NLP most often communicates progress, but in the LLM era they are increasingly easy to misread. Scores can reflect benchmark-chasing, hidden evaluation choices, or accidental exposure to test content -- not just broad capability. Closed benchmarks delay some of these issues, but reduce transparency and make it harder for the community to learn from results. We argue for a complementary practice: an Olympiad-style evaluation event where problems are sealed until evaluation, submissions are frozen in advance, and all entries run through one standardized harness. After scoring, the full task set and evaluation code are released so results can be reproduced and audited. This design aims to make strong performance harder to ``manufacture'' and easier to trust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。