用分阶段模块化方法提升AI生成教育类比喻的质量。
Teaching Through Analogies: A Modular Pipeline for Educational Analogy Generation

- 将比喻生成拆分为找原型、生成子概念、写解释、评估四步,系统分析每步影响。
- 子概念能显著提升解释质量与检索精度,但对开放源生成帮助有限。
- 提出用大模型评分,发现Claude Sonnet 4.6更贴近人类评价排序。
比喻通过将陌生概念与已知概念关联,帮助学习者理解新知识。尽管近期大语言模型(LLMs)取得进展,其生成的比喻质量仍难媲美人类。本文提出一种模块化教育比喻生成流程,将任务分解为四个阶段:原型查找、子概念生成、解释生成和评估。基于结构映射理论,该流程支持对模型选择与输入配置如何影响比喻质量进行逐阶段分析。我们在两个带有结构化子概念标注的数据集(SCAR 和 ParallelPARC)上,评估了12个顶尖LLMs(来自六个模型家族)及七种嵌入模型在封闭场景下的检索表现。结果表明,子概念显著提升了解释质量与封闭检索精度,但在开放源生成中作用有限。我们进一步引入‘大模型作为评判者’的评估方法,并与七名标注者的打分对比验证,发现Claude Sonnet 4.6在排名一致性上优于绝对分数一致性。综合来看,本研究揭示了各阶段间的相互作用,强调子概念的语义锚定是提升比喻质量的关键。
原文摘要 · Abstract (English)
Analogies help learners understand unfamiliar concepts by relating them to known concepts. Despite recent advances, large language models (LLMs) continue to struggle to generate analogies of comparable quality to those produced by humans. We present a modular pipeline for educational analogy generation, decomposing the task into four stages: source finding, sub-concept generation, explanation generation, and evaluation. Grounded in Structure Mapping Theory, the pipeline enables systematic, stage-by-stage analysis of how model choice and input configuration affect analogy quality. We evaluate 12 state-of-the-art LLMs across six model families on two datasets with structured sub-concept annotations (SCAR and ParallelPARC), alongside seven embedding models for closed-setting retrieval. Our results show that sub-concepts substantially improve explanation quality and closed setting retrieval precision but provide limited benefit in open-ended source generation. We further introduce an LLM-as-a-judge evaluation methodology and validate its scoring against human annotations from seven annotators, finding that Claude Sonnet 4.6 aligns more reliably with human rankings than with fine-grained absolute scores. Taken together, our findings reveal cross-stage interactions that isolated studies cannot capture, and highlight sub-concept grounding as a key driver of analogy quality generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。