arXiv:2601.17022cs.MMcs.AI2026-01

用AI把文字或语音转成互动教育视频,效果优于现有方法

AI-based System for Transforming text and sound to Educational Videos

  • 分三阶段:语音转文字、关键词生成图像、图像合成视频
  • FID得分28.75%,视觉质量与语义对齐优于TGAN等模型
  • 适合教育科技开发者和在线课程制作人员

技术发展已使从输入文本或声音生成教育视频成为可能。近年来,深度学习在图像与视频生成中的应用广泛探索,尤其在教育领域。然而,基于文本或语音等条件输入生成视频仍具挑战性。本文提出一种新方法,引入生成对抗网络(GAN)构建逐帧框架,实现完整教育视频生成。系统分为三个阶段:第一阶段使用语音识别将输入(文本或语音)转录;第二阶段提取关键词,并利用CLIP与扩散模型生成高质量、语义匹配的图像;第三阶段将生成图像合成视频,整合预录或合成音频,形成可交互教育视频。该系统与TGAN、MoCoGAN及TGANS-C对比,取得28.75%的弗雷谢特起始距离(FID)分数,表明视觉质量更优,性能超越现有方法。

原文摘要 · Abstract (English)

Technological developments have produced methods that can generate educational videos from input text or sound. Recently, the use of deep learning techniques for image and video generation has been widely explored, particularly in education. However, generating video content from conditional inputs such as text or speech remains a challenging area. In this paper, we introduce a novel method to the educational structure, Generative Adversarial Network (GAN), which develop frame-for-frame frameworks and are able to create full educational videos. The proposed system is structured into three main phases In the first phase, the input (either text or speech) is transcribed using speech recognition. In the second phase, key terms are extracted and relevant images are generated using advanced models such as CLIP and diffusion models to enhance visual quality and semantic alignment. In the final phase, the generated images are synthesized into a video format, integrated with either pre-recorded or synthesized sound, resulting in a fully interactive educational video. The proposed system is compared with other systems such as TGAN, MoCoGAN, and TGANS-C, achieving a Fréchet Inception Distance (FID) score of 28.75%, which indicates improved visual quality and better over existing methods.

教育AI视频生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。