arXiv:2604.02221cs.HCcs.AI2026-04中稿 · AIED 2026

多模态对话AI提升生物学习效果,仅靠对话不配图文反而误导认知。

Impact of Multimodal and Conversational AI on Learning Outcomes and Experience

  • 用图文混合对话系统教生物学,比纯文字或传统阅读更有效
  • 图文对话组成绩最高,纯文字对话组虽感觉好但实际学得差
  • 结合视觉与语言能增强专注力,适合科学教育场景

多模态大语言模型为基于教育内容的对话式多媒体学习提供了可能。然而,尽管对话式AI能提升参与度,其在视觉丰富的STEM领域对学习效果的影响仍不明确,且多模态与对话性如何协同影响生成式AI中的学习尚不清楚。本研究通过一项随机对照在线实验(N = 124)比较三种生物学学习方式:(1) 文档驱动、图文交替响应的多模态对话系统(MuDoC),(2) 文档驱动、纯文本响应的对话系统(TexDoC),(3) 带语义搜索与高亮功能的传统文本界面(DocSearch)。结果显示,使用MuDoC的学习者取得最高后测成绩,并报告最积极的学习体验。值得注意的是,尽管TexDoC在感知上比DocSearch更具吸引力且更易用,但其后测成绩最低,揭示了学习者自我感知与真实学习成果之间的脱节。根据认知负荷理论,对话性可降低外在负荷,而多模态整合则提升内在负荷,从而促进深度学习。若缺乏多模态支持,对话带来的认知简化反而会虚增理解错觉,损害实际学习效果。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) offer an opportunity to support multimedia learning through conversational systems grounded in educational content. However, while conversational AI is known to boost engagement, its impact on learning in visually-rich STEM domains remains under-explored. Moreover, there is limited understanding of how multimodality and conversationality jointly influence learning in generative AI systems. This work reports findings from a randomized controlled online study (N = 124) comparing three approaches to learning biology from textbook content: (1) a document-grounded conversational AI with interleaved text-and-image responses (MuDoC), (2) a document-grounded conversational AI with text-only responses (TexDoC), and (3) a textbook interface with semantic search and highlighting (DocSearch). Learners using MuDoC achieved the highest post-test scores and reported the most positive learning experience. Notably, while TexDoC was rated as significantly more engaging and easier to use than DocSearch, it led to the lowest post-test scores, revealing a disconnect between student perceptions and learning outcomes. Interpreted through the lens of the Cognitive Load Theory, these findings suggest that conversationality reduces extraneous load, while visual-verbal integration induced by multimodality increases germane load, leading to better learning outcomes. When conversationality is not complemented by multimodality, reduced cognitive effort may instead inflate perceived understanding without improving learning outcomes.

多模态学习对话AI教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。