首个面向中文语境的多任务多指标评测基准,支持自动评分与人类偏好融合。
BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
- 采用双评分机制,结合统计显著性与社会偏好理论进行综合评估。
- 涵盖50个本土化任务,14个新收集数据集,覆盖历史新闻、学生作文等场景。
- 开源模型与排行榜,支持持续提交,助力捷克语大模型研究发展。
我们提出BenCzechMark(BCM),首个专为捷克语大语言模型设计的综合性评测基准,包含多样任务、多种格式及多维度评价指标。其双评分系统基于统计显著性理论,并借鉴社会偏好理论进行任务聚合。基准涵盖50个具有挑战性的任务,配套测试数据集,主要为原生捷克语,其中14个为新采集。任务覆盖8个类别,涉及历史捷克新闻、学生或语言学习者作文、口语文本等多元领域。此外,我们构建并清洗了BUT-Large Czech Collection,目前最大的公开可访问清洁捷克语语料库,用于(1)数据污染分析,(2)首个基于捷克语特有分词的70亿参数捷克语中心模型的持续预训练。我们以该模型为基线,与公开多语言模型进行对比。最后,我们在Hugging Face上发布并维护排行榜,已有50个模型提交,新模型可随时提交至https://huggingface.co/spaces/CZLC/BenCzechMark。
原文摘要 · Abstract (English)
We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple evaluation metrics. Its duel scoring system is grounded in statistical significance theory and uses aggregation across tasks inspired by social preference theory. Our benchmark encompasses 50 challenging tasks, with corresponding test datasets, primarily in native Czech, with 14 newly collected ones. These tasks span 8 categories and cover diverse domains, including historical Czech news, essays from pupils or language learners, and spoken word. Furthermore, we collect and clean BUT-Large Czech Collection, the largest publicly available clean Czech language corpus, and use it for (i) contamination analysis and (ii) continuous pretraining of the first Czech-centric 7B language model with Czech-specific tokenization. We use our model as a baseline for comparison with publicly available multilingual models. Lastly, we release and maintain a leaderboard with existing 50 model submissions, where new model submissions can be made at https://huggingface.co/spaces/CZLC/BenCzechMark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。