arXiv:2502.04103cs.HCcs.AI2025-02被引 8

用生成式AI打造可动的智能教学助手,支持多模态互动。

VTutor: An Open-Source SDK for Generative AI-Powered Animated Pedagogical Agents with Multi-Media Output

  • 结合大模型与动画技术,实现实时个性化反馈
  • 支持2D/3D角色,具备自然口型同步和网页无缝集成
  • 适合教育AI开发人员,推动可信人机交互

大语言模型(LLMs)的快速发展重塑了人机交互,但当前交互仍以文本为主,多模态方式尚未充分探索。本文提出VTutor,一个开源软件开发工具包(SDK),融合生成式AI与先进动画技术,构建可互动、自适应且逼真的智能教学代理(APA),支持多模态人机交互。VTutor利用LLMs实现实时个性化反馈,采用先进唇同步技术确保语音自然对齐,并通过WebGL渲染实现无缝网页集成。支持多种2D与3D角色模型,帮助研究者与开发者设计情感共鸣、情境自适应的学习代理。该工具包提升学习者参与度、反馈接受度与人机交互体验,同时促进教育领域可信AI原则的落地。VTutor为下一代智能教学代理树立新标准,提供可访问、可扩展的解决方案,项目已开源,欢迎社区共建与展示。

原文摘要 · Abstract (English)

The rapid evolution of large language models (LLMs) has transformed human-computer interaction (HCI), but the interaction with LLMs is currently mainly focused on text-based interactions, while other multi-model approaches remain under-explored. This paper introduces VTutor, an open-source Software Development Kit (SDK) that combines generative AI with advanced animation technologies to create engaging, adaptable, and realistic APAs for human-AI multi-media interactions. VTutor leverages LLMs for real-time personalized feedback, advanced lip synchronization for natural speech alignment, and WebGL rendering for seamless web integration. Supporting various 2D and 3D character models, VTutor enables researchers and developers to design emotionally resonant, contextually adaptive learning agents. This toolkit enhances learner engagement, feedback receptivity, and human-AI interaction while promoting trustworthy AI principles in education. VTutor sets a new standard for next-generation APAs, offering an accessible, scalable solution for fostering meaningful and immersive human-AI interaction experiences. The VTutor project is open-sourced and welcomes community-driven contributions and showcases.

智能教学生成式AI多模态交互开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。