用多智能体自省机制生成图文一致的教育题目,准确率超GPT-4o三倍。
MAGMA-Edu: Multi-Agent Generative Multimodal Framework for Text-Diagram Educational Question Generation
- 分两阶段协同演化:先迭代优化题目与解法,再用代码中间表示保证图形准确性。
- 文本指标提升35.3个百分点,图文一致性达85.24,刷新教育多模态生成纪录。
- 适合教育AI研发者、智能题库构建者,特别适用于数学类可视化教学场景。
教育图示在传递抽象概念中起核心作用,但现有多模态大模型在生成具有教学逻辑和语义一致性的教育图像方面仍受限。本文提出MAGMA-Edu,一种自省式多智能体框架,统一文本推理与图示合成,实现结构化教育问题生成。不同于将文本与图像生成独立处理的方法,MAGMA-Edu采用两阶段共进化流程:(1) 生成-验证-反思循环,迭代优化数学准确性;(2) 基于代码的中间表示,在图像渲染中确保几何保真度与语义对齐。两个阶段均由内部自省模块驱动,持续评估并修正输出直至满足特定教学约束。在多模态教育基准上的实验表明,MAGMA-Edu显著优于现有先进模型。相比GPT-4o,其平均文本指标从57.01提升至92.31(+35.3),图文一致性(ITC)从13.20升至85.24(+72)。所有模型基座下,MAGMA-Edu均取得最高分(平均文本96.20,ITC 99.12),确立了多模态教育内容生成的新基准,验证了自省式多智能体协作在教学对齐的视觉-语言推理中的有效性。
原文摘要 · Abstract (English)
Educational illustrations play a central role in communicating abstract concepts, yet current multimodal large language models (MLLMs) remain limited in producing pedagogically coherent and semantically consistent educational visuals. We introduce MAGMA-Edu, a self-reflective multi-agent framework that unifies textual reasoning and diagrammatic synthesis for structured educational problem generation. Unlike existing methods that treat text and image generation independently, MAGMA-Edu employs a two-stage co-evolutionary pipeline: (1) a generation-verification-reflection loop that iteratively refines question statements and solutions for mathematical accuracy, and (2) a code-based intermediate representation that enforces geometric fidelity and semantic alignment during image rendering. Both stages are guided by internal self-reflection modules that evaluate and revise outputs until domain-specific pedagogical constraints are met. Extensive experiments on multimodal educational benchmarks demonstrate the superiority of MAGMA-Edu over state-of-the-art MLLMs. Compared to GPT-4o, MAGMA-Edu improves the average textual metric from 57.01 to 92.31 (+35.3 pp) and boosts image-text consistency (ITC) from 13.20 to 85.24 (+72 pp). Across all model backbones, MAGMA-Edu achieves the highest scores (Avg-Text 96.20, ITC 99.12), establishing a new state of the art for multimodal educational content generation and demonstrating the effectiveness of self-reflective multi-agent collaboration in pedagogically aligned vision-language reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。