arXiv:2502.19400cs.AIcs.CL2025-02ACL被引 23

用智能体生成5分钟以上的定理视频讲解,提升数学理解效果

TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding

  • 通过智能体规划生成基于Manim的长视频讲解
  • 93.8%成功率,综合评分0.77,优于传统方法
  • 适合需要深度理解定理的师生及AI教育研究者

理解领域特定定理不仅依赖文本推理,结构化视觉解释对深入理解至关重要。尽管大语言模型在文本定理推理中表现优异,但其生成连贯且具有教学意义的视觉解释能力仍面临挑战。本文提出TheoremExplainAgent,一种基于智能体的生成框架,用于制作超过5分钟的定理解释视频(使用Manim动画)。为系统评估多模态解释效果,我们构建了TheoremExplainBench基准,涵盖240个跨多个STEM学科的定理,包含5项自动化评估指标。结果表明,智能体规划对生成详细长视频至关重要;o3-mini智能体达到93.8%的成功率和0.77的综合得分。然而定量与定性分析显示,多数视频存在视觉元素布局微小问题。此外,多模态解释暴露出文本解释无法揭示的深层推理缺陷,凸显多模态解释的重要性。

原文摘要 · Abstract (English)

Understanding domain-specific theorems often requires more than just text-based reasoning; effective communication through structured visual explanations is crucial for deeper comprehension. While large language models (LLMs) demonstrate strong performance in text-based theorem reasoning, their ability to generate coherent and pedagogically meaningful visual explanations remains an open challenge. In this work, we introduce TheoremExplainAgent, an agentic approach for generating long-form theorem explanation videos (over 5 minutes) using Manim animations. To systematically evaluate multimodal theorem explanations, we propose TheoremExplainBench, a benchmark covering 240 theorems across multiple STEM disciplines, along with 5 automated evaluation metrics. Our results reveal that agentic planning is essential for generating detailed long-form videos, and the o3-mini agent achieves a success rate of 93.8% and an overall score of 0.77. However, our quantitative and qualitative studies show that most of the videos produced exhibit minor issues with visual element layout. Furthermore, multimodal explanations expose deeper reasoning flaws that text-based explanations fail to reveal, highlighting the importance of multimodal explanations.

多模态解释定理理解视频生成智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。