用AI自动生成带动态高亮的讲解视频,让静态幻灯片自动变生动。
Generating Narrated Lecture Videos from Slides with Synchronized Highlights
- 通过语音与幻灯片内容对齐,实现讲解与视觉高亮精准同步。
- 基于大模型的对齐方法在复杂内容上准确率超92%,显著优于传统方法。
- 每小时生成成本低于1美元,比人工制作节省近百倍成本。
将静态幻灯片转化为引人入胜的视频讲座需要大量时间和精力,需演讲者录制解说并引导观众关注重点。本文提出一个端到端系统,可完全自动化该过程。给定幻灯片文稿,系统生成包含AI语音解说和动态视觉高亮的视频讲座,高亮位置自动匹配讲解内容,如同优秀讲师般引导观众注意力。核心技术是新颖的高亮对齐模块,通过多种策略(如莱文斯坦距离、基于大模型的语义分析)在行或词级别实现语音与幻灯片位置的精确映射,并利用提供时间戳的文本转语音(TTS)实现同步。在包含1000个样本的手动标注数据集上的技术评估显示,基于大模型的对齐方法在位置识别上取得超过92%的F1值,尤其在数学密集型内容中表现远超简单方法。此外,平均生成成本低于每小时1美元,相比保守估算的人工制作成本降低两个数量级。该方法兼具高精度与极低成本,具备实际应用与大规模推广潜力。
原文摘要 · Abstract (English)
Turning static slides into engaging video lectures takes considerable time and effort, requiring presenters to record explanations and visually guide their audience through the material. We introduce an end-to-end system designed to automate this process entirely. Given a slide deck, this system synthesizes a video lecture featuring AI-generated narration synchronized precisely with dynamic visual highlights. These highlights automatically draw attention to the specific concept being discussed, much like an effective presenter would. The core technical contribution is a novel highlight alignment module. This module accurately maps spoken phrases to locations on a given slide using diverse strategies (e.g., Levenshtein distance, LLM-based semantic analysis) at selectable granularities (line or word level) and utilizes timestamp-providing Text-to-Speech (TTS) for timing synchronization. We demonstrate the system's effectiveness through a technical evaluation using a manually annotated slide dataset with 1000 samples, finding that LLM-based alignment achieves high location accuracy (F1 > 92%), significantly outperforming simpler methods, especially on complex, math-heavy content. Furthermore, the calculated generation cost averages under $1 per hour of video, offering potential savings of two orders of magnitude compared to conservative estimates of manual production costs. This combination of high accuracy and extremely low cost positions this approach as a practical and scalable tool for transforming static slides into effective, visually-guided video lectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。