将视频讲座转化为知识驱动的可视化摘要,降低学习负担。
KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

- 基于多模态内容提取概念图,识别关键难点概念
- 构建结构化知识单元,生成清晰视觉摘要
- 适合教育科技、智能教学系统开发者使用
视频讲座是重要的教育资源,但其信息密集且冗长,常使初学者难以应对。根本原因在于:视频线性传递瞬时信息,而人类学习需构建互联认知网络,缺乏领域知识的新手易产生严重认知负荷。现有摘要方法多生成文字密集的线性压缩内容,仍需高认知努力。为此,我们提出KnowVis框架,将线性视频讲座转化为基于教学原理的视觉叙事。该框架首先从多模态视频内容中提取详细概念图,识别重要且具挑战性的门槛概念;随后构建结构化知识单元;最后合成吸引人的视觉摘要。我们还构建了一个包含125个跨10个学科的教育视频数据集,配套1,079个生成的视觉摘要。大量自动评估与人工研究表明,相比先进基线,KnowVis生成的视觉更准确、更清晰,有效降低认知负荷,显著提升学生学习效果与知识保留率。
原文摘要 · Abstract (English)
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe cognitive overload for novice learners lacking prior domain knowledge. Existing video summarization methods fail to resolve this mismatch, as they primarily produce text-heavy, linear condensations that still demand high cognitive effort. To bridge this gap, we propose KnowVis, a framework that transforms linear video lectures into pedagogically grounded visual narratives. KnowVis first extracts a detailed concept map from multimodal video content to identify important and challenging threshold concepts, then constructs structured knowledge units, and finally synthesizes engaging visual summaries. Alongside the framework, we introduce a curated dataset of 125 educational videos across 10 academic disciplines, paired with 1,079 generated visual summaries. Extensive automated evaluations and a human study demonstrate that, compared to state-of-the-art baselines, KnowVis generates more accurate and clear visuals that successfully reduce cognitive load and significantly improve student learning effectiveness and knowledge retention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。