arXiv:2504.03197cs.CL2025-04AAAI被引 3

让AI像导师一样用关键点解释数学题,提升学习理解效果

Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation

  • 设计新任务:让模型识别图形中的关键点并结合解释
  • 构建1000道带视觉标注的数学题数据集,评估模型表现
  • 发现当前模型在视觉定位和解释上仍有明显短板

随着大语言模型在数学推理能力上的快速发展,其在教育场景中辅助学生理解解题过程的应用日益广泛。然而,当前模型生成的解释仍缺乏多模态特性——真实教学中导师常借助图示、标记和高亮等视觉手段增强理解。为此,我们提出多模态解题解释任务,旨在评估模型是否能识别辅助线、点、角等视觉关键点,并生成包含这些要素的解释。为评估该任务,我们构建了ME2基准数据集,包含1000个数学问题,每个均标注了视觉关键点及对应引用这些元素的解释文本。实验表明,现有模型在识别视觉关键点方面表现不佳,开源模型在基于关键点的解释生成中也面临显著挑战。这揭示了当前大语言模型在数学视觉定位、视觉化推理以及教育场景解释能力上的明显不足。我们期望该任务与数据集能推动教育领域大模型研究,促进其作为可解释性智能导师的应用。

原文摘要 · Abstract (English)

With the rapid advancement of mathematical reasoning capabilities in Large Language Models (LLMs), AI systems are increasingly being adopted in educational settings to support students' comprehension of problem-solving processes. However, a critical component remains underexplored in current LLM-generated explanations: multimodal explanation. In real-world instructional contexts, human tutors routinely employ visual aids, such as diagrams, markings, and highlights, to enhance conceptual clarity. To bridge this gap, we introduce the multimodal solution explanation task, designed to evaluate whether models can identify visual keypoints, such as auxiliary lines, points, angles, and generate explanations that incorporate these key elements essential for understanding. To evaluate model performance on this task, we propose ME2, a multimodal benchmark consisting of 1,000 math problems annotated with visual keypoints and corresponding explanatory text that references those elements. Our empirical results show that current models struggle to identify visual keypoints. In the task of generating keypoint-based explanations, open-source models also face notable difficulties. This highlights a significant gap in current LLMs' ability to perform mathematical visual grounding, engage in visually grounded reasoning, and provide explanations in educational contexts. We expect that the multimodal solution explanation task and the ME2 dataset will catalyze further research on LLMs in education and promote their use as effective, explanation-oriented AI tutors.

多模态解释数学推理教育AI视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。