评测大模型解俄语数理难题能力的开源基准
DOoM: Difficult Olympiads of Math
- 构建涵盖中小学到奥赛级的俄语数理题集
- 发现模型表现与上下文长度正相关
- 揭示数学题比物理题更易被模型解决
本文提出DOoM,一个用于评估语言模型在俄语环境下解决数学与物理问题能力的开源基准。该基准包含从校级作业到大学奥赛及入学考试题目的多难度题目。文中阐述了创建动机,描述了数据集结构与评估方法,并展示了多种模型的初步测试结果。分析表明,模型性能与使用令牌数呈正相关,且数学任务表现普遍优于物理任务。
原文摘要 · Abstract (English)
This paper introduces DOoM, a new open-source benchmark designed to assess the capabilities of language models in solving mathematics and physics problems in Russian. The benchmark includes problems of varying difficulty, ranging from school-level tasks to university Olympiad and entrance exam questions. In this paper we discuss the motivation behind its creation, describe dataset's structure and evaluation methodology, and present initial results from testing various models. Analysis of the results shows a correlation between model performance and the number of tokens used, and highlights differences in performance between mathematics and physics tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。