arXiv:2603.25870cs.CVcs.LG2026-03

用视觉语言模型让手绘白板与语音同步,首次实现毫秒级精准绘制。

Speech-Synchronized Whiteboard Generation via VLM-Driven Structured Drawing Representations

  • 基于视觉语言模型,从少量演示中学习语音驱动的结构化绘图序列。
  • 在8个STEM领域上实现毫秒级时间对齐,跨主题泛化能力良好。
  • 适合教育视频自动化生成研究者,数据集开源可复现。

生成白板风格教育视频需精确协调手绘图与语音解说,但现有方法缺乏结构化、可复现的绘图表示来解决多模态同步问题。本文构建首个包含24组配对的Excalidraw演示数据集,涵盖8个STEM领域,每个绘图元素带有毫秒级创建时间戳。利用该数据,我们研究仅通过24个演示样本,经LoRA微调的Qwen2-VL-7B视觉语言模型能否从语音预测完整笔画序列。五折分主题评估表明,时间戳条件显著提升时间对齐性能,模型在未见主题间具有良好泛化能力。论文讨论了向真实课堂场景迁移的可行性,并公开数据集与代码,以推动自动教育内容生成研究。

原文摘要 · Abstract (English)

Creating whiteboard-style educational videos demands precise coordination between freehand illustrations and spoken narration, yet no existing method addresses this multimodal synchronization problem with structured, reproducible drawing representations. We present the first dataset of 24 paired Excalidraw demonstrations with narrated audio, where every drawing element carries millisecond-precision creation timestamps spanning 8 STEM domains. Using this data, we study whether a vision-language model (Qwen2-VL-7B), fine-tuned via LoRA, can predict full stroke sequences synchronized to speech from only 24 demonstrations. Our topic-stratified five-fold evaluation reveals that timestamp conditioning significantly improves temporal alignment over ablated baselines, while the model generalizes across unseen STEM topics. We discuss transferability to real classroom settings and release our dataset and code to support future research in automated educational content generation.

白板生成多模态同步视觉语言模型教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。