用动画视频解释物理题,让复杂公式变直观。
PhysicsSolutionAgent: Towards Multimodal Explanations for Numerical Physics Problem Solving
- 自动生成6分钟物理题动画视频,结合Manim与LLM推理
- 平均自动评分3.8/5,100%完成率但存在视觉错误
- 适合教育科技研究者和需要可视化教学的教师
解释数值物理问题常需超越文字解答,清晰的视觉推理可显著提升理解。尽管大语言模型在文本形式的物理题上表现良好,但其生成高质量长视频的能力仍待探索。本文提出PhysicsSolutionAgent(PSA),一个能生成长达六分钟物理问题讲解视频的自主代理,使用Manim动画实现。为评估视频质量,设计包含15个量化参数的自动化评估流程,并引入视觉-语言模型反馈以迭代优化。在32个涵盖数值与理论物理的问题上测试,结果表明视频质量随题目难度及类型不同而有系统性差异。使用GPT-5-mini时,视频完成率达100%,平均自动评分3.8/5。但定性分析与人工检查发现存在布局不一致及内容误读等重大问题,暴露了可靠Manim代码生成的局限,也凸显了多模态推理与评估在物理可视化中的挑战。本工作强调未来多模态教育系统需加强视觉理解、验证与评估框架。
原文摘要 · Abstract (English)
Explaining numerical physics problems often requires more than text-based solutions; clear visual reasoning can substantially improve conceptual understanding. While large language models (LLMs) demonstrate strong performance on many physics questions in textual form, their ability to generate long, high-quality visual explanations remains insufficiently explored. In this work, we introduce PhysicsSolutionAgent (PSA), an autonomous agent that generates physics-problem explanation videos of up to six minutes using Manim animations. To evaluate the generated videos, we design an assessment pipeline that performs automated checks across 15 quantitative parameters and incorporates feedback from a vision-language model (VLM) to iteratively improve video quality. We evaluate PSA on 32 videos spanning numerical and theoretical physics problems. Our results reveal systematic differences in video quality depending on problem difficulty and whether the task is numerical or theoretical. Using GPT-5-mini, PSA achieves a 100% video-completion rate with an average automated score of 3.8/5. However, qualitative analysis and human inspection uncover both minor and major issues, including visual layout inconsistencies and errors in how visual content is interpreted during feedback. These findings expose key limitations in reliable Manim code generation and highlight broader challenges in multimodal reasoning and evaluation for visual explanations of numerical physics problems. Our work underscores the need for improved visual understanding, verification, and evaluation frameworks in future multimodal educational systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。