arXiv:2512.20858cs.CV2025-12

让录播课变互动课堂,本地运行实时答疑。

ALIVE: An Avatar-Lecture Interactive Video Engine with Content-Aware Retrieval for Real-Time Interaction

  • 用语音识别+大模型优化生成虚拟讲师讲解。
  • 结合语义与时间戳检索,精准定位相关课程片段。
  • 支持本地运行,适合隐私敏感的医学教学场景。

传统录播课虽灵活但无法实时答疑,学习者困惑时需外部搜索。现有交互系统多依赖云端、缺乏课程感知或难以统一整合检索与虚拟讲师回应。我们提出ALIVE——一种全本地部署的虚拟讲师互动视频引擎,将被动听课变为实时互动学习。系统通过语音识别转写、大模型润色及神经头像合成生成虚拟讲师内容;采用语义相似性与时间对齐相结合的内容感知检索机制,精准定位上下文相关课程片段;支持学生暂停、文本或语音提问,并以文字或虚拟讲师形式获得有依据的即时反馈。为保障响应速度,系统使用轻量嵌入模型、基于FAISS的检索和分段预加载的头像合成。我们在完整医学影像课程上验证了系统,评估其检索准确率、延迟表现与用户体验,结果表明ALIVE能提供准确、情境感知且富有吸引力的实时学习支持。该研究展示,融合内容感知检索与本地化部署的多模态AI可显著提升录播课的教学价值,为下一代互动学习环境提供可扩展路径。

原文摘要 · Abstract (English)

Traditional lecture videos offer flexibility but lack mechanisms for real-time clarification, forcing learners to search externally when confusion arises. Recent advances in large language models and neural avatars provide new opportunities for interactive learning, yet existing systems typically lack lecture awareness, rely on cloud-based services, or fail to integrate retrieval and avatar-delivered explanations in a unified, privacy-preserving pipeline. We present ALIVE, an Avatar-Lecture Interactive Video Engine that transforms passive lecture viewing into a dynamic, real-time learning experience. ALIVE operates fully on local hardware and integrates (1) Avatar-delivered lecture generated through ASR transcription, LLM refinement, and neural talking-head synthesis; (2) A content-aware retrieval mechanism that combines semantic similarity with timestamp alignment to surface contextually relevant lecture segments; and (3) Real-time multimodal interaction, enabling students to pause the lecture, ask questions through text or voice, and receive grounded explanations either as text or as avatar-delivered responses. To maintain responsiveness, ALIVE employs lightweight embedding models, FAISS-based retrieval, and segmented avatar synthesis with progressive preloading. We demonstrate the system on a complete medical imaging course, evaluate its retrieval accuracy, latency characteristics, and user experience, and show that ALIVE provides accurate, content-aware, and engaging real-time support. ALIVE illustrates how multimodal AI-when combined with content-aware retrieval and local deployment-can significantly enhance the pedagogical value of recorded lectures, offering an extensible pathway toward next-generation interactive learning environments.

互动教学虚拟讲师本地部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。