用多模态证据构建可追溯的教育知识图谱,提升讲座理解准确性。
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning

- 融合语音、字幕、图像和视觉信息,仅提取有证据支持的概念与关系。
- 在3个神经网络讲座上处理3118帧,保留1022个概念和312个关系,覆盖90.38%关键节点。
- 适合需要可解释、可验证知识系统的教育研究者和智能教学系统开发者。
讲座视频中的知识分散在语音、幻灯片文字、图表、公式和演示顺序中,仅依赖字幕的检索无法完整保留。本文提出一种基于证据的多模态流程:对讲座进行语音转写,选取语义锚点,使用光学字符识别(OCR)提取文字,并通过视觉-语言模型仅提取由字幕、OCR或视觉证据支持的概念与类型化关系。所有提及项经验证并规范化为带来源信息的知识图谱。在三个神经网络讲座上,该流程处理了3,118帧、756段字幕和559个锚点,保留了1,022个概念提及和312个关系提及,生成172个标准化概念和282个关系,终点覆盖率高达90.38%。初步的三题检索测试达到100%的top-1和top-3准确率,以及100%的平均top-5召回率。贡献在于可审计的构建方法,而非性能突破。
原文摘要 · Abstract (English)
Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve. This paper presents an evidence-grounded multimodal pipeline that transcribes lectures, selects semantic anchors, applies optical character recognition (OCR), and uses a vision-language model to extract only concepts and typed relationships supported by transcript, OCR, or visual evidence. Mentions are validated and canonicalized into a provenance-rich knowledge graph. On three neural-network lectures, the pipeline processed 3,118 frames, 756 transcript segments, and 559 anchors. It retained 1,022 concept and 312 relationship mentions, yielding 172 canonical concepts and 282 relationships with 90.38% endpoint coverage. A preliminary three question retrieval test achieved 100% top-1 and top-3 accuracy and 100% mean top-5 recall. The contribution is an auditable construction method rather than a state-of-the-art performance claim.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。