用中间表示自动评估数学图示,提升教育AI的可视化能力
DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
- 将LaTeX TikZ代码转为中间表示,实现可扩展的图示评估
- 相比人工评分,评估一致性更高,且小模型成本降低10倍
- 适合教育AI、智能题库开发人员使用
大型语言模型(LLM)在学习工具中的应用日益广泛,但多数工具仍局限于文本,难以满足数学等需要视觉表达的领域。已有研究证明LLM能生成可编译为教学图形的代码,但大规模评估仍是主要瓶颈。为此,本文提出DiagramIR:一种基于LaTeX TikZ代码中间表示(IR)的自动化、可扩展几何图形评估流水线。与LLM作为裁判等基线方法相比,该方法在与人类评分者的一致性上表现更优。此外,该评估方式使GPT-4.1-Mini等小型模型在推理成本降低10倍的情况下,性能可媲美GPT-5等大型模型,对部署可及、可扩展的教育技术具有重要意义。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being adopted as tools for learning; however, most tools remain text-only, limiting their usefulness for domains where visualizations are essential, such as mathematics. Recent work shows that LLMs are capable of generating code that compiles to educational figures, but a major bottleneck remains: scalable evaluation of these diagrams. We address this by proposing DiagramIR: an automatic and scalable evaluation pipeline for geometric figures. Our method relies on intermediate representations (IRs) of LaTeX TikZ code. We compare our pipeline to other evaluation baselines such as LLM-as-a-Judge, showing that our approach has higher agreement with human raters. This evaluation approach also enables smaller models like GPT-4.1-Mini to perform comparably to larger models such as GPT-5 at a 10x lower inference cost, which is important for deploying accessible and scalable education technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。