构建无污染双语评估基准,让大模型评测更可信。
C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation
- 设计22项任务覆盖大模型多种能力,全面评估性能。
- 自动化更新测试数据并保护隐私,确保无训练数据泄露。
- 适合关注模型真实能力、需公平对比的开发者和研究者。
大语言模型虽有显著进展,但评估面临数据污染问题,因无法获取专有训练数据。为此,我们提出C$^2$LEVA,一个全面且无污染的双语评估基准。该基准包含22个任务,涵盖大模型的各项应用与能力;通过系统化的污染预防策略,完全自动化测试数据更新,并在发布时强制数据保护,确保评估可信。对15个开源与专有模型的大规模评估表明,C$^2$LEVA能有效衡量模型真实性能。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have shown significant promise, yet their evaluation raises concerns, particularly regarding data contamination due to the lack of access to proprietary training data. To address this issue, we present C$^2$LEVA, a comprehensive bilingual benchmark featuring systematic contamination prevention. C$^2$LEVA firstly offers a holistic evaluation encompassing 22 tasks, each targeting a specific application or ability of LLMs, and secondly a trustworthy assessment due to our contamination-free tasks, ensured by a systematic contamination prevention strategy that fully automates test data renewal and enforces data protection during benchmark data release. Our large-scale evaluation of 15 open-source and proprietary models demonstrates the effectiveness of C$^2$LEVA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。