arXiv:2504.13884cs.HCcs.AI2025-04被引 9

让AI教书时能用图文互动,并可查证来源,提升学生信任感。

Towards a Multimodal Document-grounded Conversational AI System for Education

  • 用图文混合方式生成回答,支持跨模态内容理解与输出。
  • 加入可追溯来源功能后,学生参与度和对AI的信任显著提升。
  • 适合教育类AI系统研发者,关注人机交互可信性与学习效果。

多媒体学习结合文本与图像相比纯文本教学更有利于提升学习效果。然而,当前教育领域的对话式AI系统仍以纯文本交互为主,多模态对话在多媒体学习中的应用尚未被探索。此外,教育场景中部署对话AI需具备可靠知识来源和内容可验证性以建立信任。本文提出MuDoC——一个基于GPT-4o的多模态文档接地对话AI系统,能够利用文档中的文本与视觉信息生成包含图文的响应,并通过界面实现对AI输出内容的无缝溯源。我们对比了MuDoC与纯文本系统在学习者参与度、对AI的信任度以及问题解决表现上的差异。结果表明,图文结合与内容可验证性均能显著提升学习者参与度与信任感,但对任务表现无显著影响。研究基于认知与学习科学理论解释发现,并提出未来多模态教育对话AI的发展方向。

原文摘要 · Abstract (English)

Multimedia learning using text and images has been shown to improve learning outcomes compared to text-only instruction. But conversational AI systems in education predominantly rely on text-based interactions while multimodal conversations for multimedia learning remain unexplored. Moreover, deploying conversational AI in learning contexts requires grounding in reliable sources and verifiability to create trust. We present MuDoC, a Multimodal Document-grounded Conversational AI system based on GPT-4o, that leverages both text and visuals from documents to generate responses interleaved with text and images. Its interface allows verification of AI generated content through seamless navigation to the source. We compare MuDoC to a text-only system to explore differences in learner engagement, trust in AI system, and their performance on problem-solving tasks. Our findings indicate that both visuals and verifiability of content enhance learner engagement and foster trust; however, no significant impact in performance was observed. We draw upon theories from cognitive and learning sciences to interpret the findings and derive implications, and outline future directions for the development of multimodal conversational AI systems in education.

多模态教育AI可信生成图文交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。