用真人协作框架评估大模型在中学数学能力测评中的表现。
Human-in-the-Loop Benchmarking of Heterogeneous LLMs for Automated Competency Assessment in Secondary Level Mathematics
- 设计多维评分标准,结合真实教师标注验证多个大模型表现。
- 70B参数的Orion模型与教师判断几乎无一致性(kappa=-0.026)。
- 模型架构适配性比参数规模更重要,适合教育辅助而非独立认证。
随着基于能力的教育(CBE)在全球兴起,从分数评价转向定性能力映射对教师而言仍是人工挑战。本文提出一种‘真人协作’基准测试框架,评估多种大模型在中学数学自动化测评中的有效性。基于尼泊尔十年级选修数学课程,构建涵盖四个主题和四项跨领域能力(理解、知识、操作熟练度、行为与关联)的多维评分体系。对比由两位资深数学教师定义的基准(kappa_w = 0.8652),测试了包含开源模型Eagle(Llama 3.1-8B)与Orion(Llama 3.3-70B),以及专有前沿模型Nova(Gemini 2.5 Flash)与Lyra(Gemini 3 Pro)的多供应商集成系统。结果表明存在显著的‘架构兼容性差距’:尽管基于Gemini的稀疏专家混合(Sparse MoE)模型达到‘尚可一致’(kappa_w ~ 0.38),但更大规模的Orion(70B)模型却出现‘无一致’(kappa_w = -0.0261),说明在评分约束任务中,架构对指令要求的适配性优于参数总量。结论指出,当前大模型尚不适合自主认证,但在‘真人协作’框架下可为初步证据提取提供高价值辅助。
原文摘要 · Abstract (English)
As Competency-Based Education (CBE) is gaining traction around the world, the shift from marks-based assessment to qualitative competency mapping is a manual challenge for educators. This paper tackles the bottleneck issue by suggesting a "Human-in-the-Loop" benchmarking framework to assess the effectiveness of multiple LLMs in automating secondary-level mathematics assessment. Based on the Grade 10 Optional Mathematics curriculum in Nepal, we created a multi-dimensional rubric for four topics and four cross-cutting competencies: Comprehension, Knowledge, Operational Fluency, and Behavior and Correlation. The multi-provider ensemble, consisted of open-weight models -- Eagle (Llama 3.1-8B) and Orion (Llama 3.3-70B) -- and proprietary frontier models Nova (Gemini 2.5 Flash) and Lyra (Gemini 3 Pro), was benchmarked against a ground truth defined by two senior mathematics faculty members (kappa_w = 0.8652). The findings show a marked "Architecture-compatibility gap". Although the Gemini-based Mixture-of-Experts (Sparse MoE) models achieved "Fair Agreement" (kappa_w ~ 0.38), the larger Orion (70B) model exhibited "No Agreement" (kappa_w = -0.0261), suggesting that architectural compliance with instruction constraints outweighs the scale of raw parameters in rubric-constrained tasks. We conclude that while LLMs are not yet suitable for autonomous certification, they provide high-value assistive support for preliminary evidence extraction within a "Human-in-the-Loop" framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。