arXiv:2605.31212cs.CVcs.AI2026-05

让AI根据算术题生成符合教学逻辑的图像,提升教育内容准确性。

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

论文配图:Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
图 1 · 摘自论文原文
  • 从算术题生成教学图像,强调数值与关系结构准确
  • 现有模型常错在物体数量和关系表达上,错误率高
  • 适合教育科技开发者和关注精准视觉生成的研究者

人工智能系统越来越多地用于支持教育内容创作,但尚不清楚它们能否生成忠实反映教学概念的输出。为此,我们提出方程到视觉生成任务,该任务要求从算术方程生成具有教学意义的图像,并精确保持其数值和关系结构。基于对教师的访谈和对教育材料的分析,我们构建了E2V-Bench基准,涵盖四种教学基础的视觉类型,并设计了自动评估指标。评估发现,当前主流文本到图像(T2I)模型在此任务中表现不佳,主要错误为物体数量错误和关系结构断裂。在此基础上,我们探索了基准引导的增强策略,有效提升了代表性模型性能,但仍存在差距,表明未来T2I模型需更强的数值与关系建模能力。

原文摘要 · Abstract (English)

AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully represent the pedagogical concepts they are intended to teach. Thus, we introduce equation-to-visual generation, a task that, in contrast to conventional image generation, requires producing pedagogically meaningful visuals from arithmetic equations while precisely preserving their numerical and relational structure. Informed by interviews with teachers and an analysis of educational materials, we construct E2V-Bench, a benchmark spanning four pedagogically grounded visual types, along with automatic metrics for evaluating visual correctness. Our evaluation reveals that recent text-to-image (T2I) models frequently fail on this task, with errors dominated by incorrect object counts and broken relational structure. Building on this, we explore benchmark-guided enhancement strategies. These strategies improve representative models, while the remaining gap calls for stronger numerical and relational grounding in future T2I models.

教育AI图像生成数学教学视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。