arXiv:2505.18549cs.CL2025-05被引 6

用一致性感知方法提升大模型数学辅导评估的多维度准确性

MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors

  • 统一微调框架,不改结构跨四个维度评估
  • 引入分歧感知集成策略,提升少数类标签覆盖
  • 在指导性上排名第一,适合教育AI评估研究者

我们提出MSA-MathEval,参加BEA 2025共享任务中对AI辅导回答在四个教学维度上的评估:错误识别、错误定位、提供指导和可操作性。我们的方法采用统一训练流程,在不进行任何任务特异性架构调整的前提下,对单一指令微调语言模型进行全维度微调。为提高预测可靠性,引入分歧感知集成推理策略,增强对少数标签的覆盖能力。系统在所有赛道表现优异,其中在提供指导维度排名1st,可操作性排名第3,错误识别与错误定位均排名第4。结果表明,可扩展的指令微调与分歧驱动建模在构建稳健的多维大模型教育评估体系中具有显著有效性。

原文摘要 · Abstract (English)

We present MSA-MathEval, our submission to the BEA 2025 Shared Task on evaluating AI tutor responses across four instructional dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability. Our approach uses a unified training pipeline to fine-tune a single instruction-tuned language model across all tracks, without any task-specific architectural changes. To improve prediction reliability, we introduce a disagreement-aware ensemble inference strategy that enhances coverage of minority labels. Our system achieves strong performance across all tracks, ranking 1st in Providing Guidance, 3rd in Actionability, and 4th in both Mistake Identification and Mistake Location. These results demonstrate the effectiveness of scalable instruction tuning and disagreement-driven modeling for robust, multi-dimensional evaluation of LLMs as educational tutors.

大模型评估教育AI多维度评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。