arXiv:2506.03735cs.CLcs.AI2025-06ACL被引 8

自动生成符合教学逻辑的数学题配图,提升学习效果。

Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models

  • 基于教师访谈设计可视化语言和布局空间,精准表达数学关系。
  • 构建1903张标注配图数据集,评估并优化文本生成图像模型。
  • 解决图像错配数学关系、遗漏关键元素等教育视觉生成难题。

可视化是教授数学应用题(MWPs)的有效工具,帮助年幼儿童将文字描述转化为数学表达式。然而,手工制作这类视觉材料耗时费力,缺乏自动化支持方法。本文提出Math2Visual,一个从数学题文本自动生成具有教学意义配图的框架。该框架基于与数学教师访谈所得的预定义视觉语言及设计空间,准确呈现题中核心数学关系。利用Math2Visual,我们构建了包含1,903张标注图像的数据集,并评估了多种文本到图像(TTI)模型在生成符合教学设计图像方面的能力。进一步用该数据集微调多个TTI模型,显著提升了教育类图像生成效果。本研究建立了自动化生成教学有意义视觉内容的新基准,揭示了多模态教育内容生成中的关键挑战,如数学关系误表示、关键视觉元素缺失等问题。

原文摘要 · Abstract (English)

Visuals are valuable tools for teaching math word problems (MWPs), helping young learners interpret textual descriptions into mathematical expressions before solving them. However, creating such visuals is labor-intensive and there is a lack of automated methods to support this process. In this paper, we present Math2Visual, an automatic framework for generating pedagogically meaningful visuals from MWP text descriptions. Math2Visual leverages a pre-defined visual language and a design space grounded in interviews with math teachers, to illustrate the core mathematical relationships in MWPs. Using Math2Visual, we construct an annotated dataset of 1,903 visuals and evaluate Text-to-Image (TTI) models for their ability to generate visuals that align with our design. We further fine-tune several TTI models with our dataset, demonstrating improvements in educational visual generation. Our work establishes a new benchmark for automated generation of pedagogically meaningful visuals and offers insights into key challenges in producing multimodal educational content, such as the misrepresentation of mathematical relationships and the omission of essential visual elements.

数学教育图文生成教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。