用视觉语言检索增强生成,让自动教学更懂学生需求。
Automatic Teaching Platform on Vision Language Retrieval Augmented Generation
- 基于视觉与语言的检索增强生成机制,动态匹配问题与答案
- 通过图文结合的问答提升理解深度,减少对人工干预依赖
- 适合需要个性化辅导的在线教育场景,可拓展至多学科
自动化教学面临挑战,因难以复现人类互动与适应性。现有系统常无法提供符合学生个体学习节奏与理解水平的细腻实时反馈,尤其在抽象概念教学中更为明显。本文提出一种视觉语言检索增强生成(VL-RAG)系统,通过调用定制化答案与图像数据库,动态检索与问题相关的上下文信息,生成兼具视觉与语言内容的响应,从而提升理解效果。该系统支持学生以图文方式探索知识,增强参与感与主动性,减少对持续人工监督的需求,并具备跨学科扩展能力。
原文摘要 · Abstract (English)
Automating teaching presents unique challenges, as replicating human interaction and adaptability is complex. Automated systems cannot often provide nuanced, real-time feedback that aligns with students' individual learning paces or comprehension levels, which can hinder effective support for diverse needs. This is especially challenging in fields where abstract concepts require adaptive explanations. In this paper, we propose a vision language retrieval augmented generation (named VL-RAG) system that has the potential to bridge this gap by delivering contextually relevant, visually enriched responses that can enhance comprehension. By leveraging a database of tailored answers and images, the VL-RAG system can dynamically retrieve information aligned with specific questions, creating a more interactive and engaging experience that fosters deeper understanding and active student participation. It allows students to explore concepts visually and verbally, promoting deeper understanding and reducing the need for constant human oversight while maintaining flexibility to expand across different subjects and course material.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。