构建首个几何题解多模态对话数据集,支持视觉标注教学。
GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns

- 通过教师实录构建1300+组含图示标注的师生对话
- 视觉标注生成准确率仅42%,暴露当前模型缺陷
- 适合研究视觉辅助教学、多模态教育AI的学者
许多教育领域高度依赖图表和视觉线索,但现有辅导数据集多为纯文本交互,限制了能进行视觉化教学的AI导师发展。为此,我们提出GeoDial,一个来自经验数学教师的超过1300组几何问题求解的多模态辅导对话数据集,其中教学语句明确基于图示高亮。我们设计了一种可扩展的标注协议,整合对话行为、视觉高亮与反馈,实现语言与视觉教学行为的细粒度监督。为展示该场景的挑战,我们在GeoDial上微调多个视觉-语言模型,并评估其生成辅导语句和图示高亮的能力。尽管监督微调显著提升了对话质量,但在生成准确图示高亮方面仍表现不佳(准确率仅42%),揭示了当前方法在视觉推理与教学互动融合上的关键局限,凸显了对更优融合策略的需求。
原文摘要 · Abstract (English)
Several educational domains rely heavily on diagrams and visual cues, yet most existing tutoring datasets are limited to text-only interactions. This limits the development of AI tutors that can teach in visually grounded ways used by human instructors. Thus, we introduce GeoDial, a multimodal tutoring dataset of over 1.3K teacher-student dialogs in the domain of geometry collected from experienced math teachers, where instructional turns are explicitly grounded in diagram highlights. We propose a scalable annotation protocol that integrates dialog acts, visual highlighting, and feedback, enabling fine-grained supervision of both language and visual tutoring behavior. To illustrate the challenges posed by this setting, we fine-tune several vision-language models on GeoDial and evaluate their ability to generate tutoring utterances and diagram highlights. While supervised fine-tuning substantially improves the quality of generated dialog, it struggles to produce accurate diagram highlights, revealing a key limitation of current methods and highlighting the need for approaches that more effectively integrate visual reasoning with pedagogical interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。