arXiv:2604.16392cs.CYcs.AI2026-04

罗马尼亚数学考试数据集,覆盖1895-2025年,可用于教育AI研究。

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

论文配图:RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)
图 1 · 摘自论文原文
  • 构建了1895-2025年罗马尼亚中学数学考试的长期数据集,核心为1957-2025年标准化题库。
  • 包含10,592道题目,覆盖600多个完整试卷,支持难度评估与相似性检索。
  • 适合教育研究、课程分析及低资源语言下大模型评估,具有可复现性。

教育领域中的人工智能研究越来越依赖真实且与课程相关的评估数据,但许多语言和教育体系仍缺乏大型、结构良好的考试语料库。本文提出 RoMathExam,一个涵盖1895–2025年的罗马尼亚高中数学考试纵向数据集,其中1957–2025年部分具有标准化核心。数据集包含10,592道数学题,分为600多个完整试卷,覆盖M1-M4多个考试类别,包含官方国考及教育部发布的训练题型。除高保真数字化与统一的JSON格式外,还添加了课程对齐的主题标签和密集文本嵌入,支持题型变体检测、去重与基于相似性的检索。为弥补历史心理测量数据缺失,我们提出并验证了一种可扩展的解题复杂度指标,作为难度的内在代理。在三款前沿推理模型(GPT-5-mini、DeepSeek-R1、Qwen3-235B-Thinking)上的评估显示,跨模型同步性高于0.72,证明该指标能有效分离内在数学深度与生成噪声。通过纵向分析,我们量化出从多变的历史题型向标准化、代数主导的现代课程的“制度转型”。RoMathExam为难度建模、课程分析及低资源语境下的大模型评估提供了可复现的研究基础。

原文摘要 · Abstract (English)

AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce for many languages and educational systems. We introduce RoMathExam, a longitudinal dataset of Romanian high-school mathematics exams spanning 1895-2025, with a robust standardized core for 1957-2025. The dataset contains 10,592 mathematics problems organized into 600+ complete exam sets across multiple tracks (M1-M4), covering both official national examination sessions and ministry-published training variants. Beyond high-fidelity digitization and a unified JSON schema with traceable provenance, RoMathExam is enriched with curriculum-aligned topic tags and dense text embeddings, enabling variant detection, deduplication, and similarity-based retrieval. To overcome the lack of historical psychometric data, we propose and validate a solution complexity metric as a scalable intrinsic proxy for difficulty. Our evaluation across three frontier reasoning models (GPT-5-mini, DeepSeek-R1, and Qwen3-235B-Thinking) reveals high cross-model synchronization (r > 0.72), confirming the metric's ability to isolate intrinsic mathematical depth from stochastic generation noise. We demonstrate the dataset's utility through a longitudinal analysis that quantifies a "regime shift" from volatile historical formats to a standardized, algebra-dominant modern curriculum. RoMathExam provides a foundation for reproducible research in difficulty modeling, curriculum analytics, and LLM evaluation in low-resource linguistic contexts.

教育AI数学考试数据集课程分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。