构建分层知识图谱,让机器读懂漫画的叙事逻辑。
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
- 按画面、事件、宏观情节三层建模,融合语义时空关系
- 在Manga109数据集上完成动作检索等四类任务
- 强调可解释性,契合人类对故事的感知机制
我们提出一种用于视觉叙事结构化语义理解的分层知识图谱框架,以漫画为多模态叙事代表。该框架在画面、事件和宏观事件三个层级组织叙事内容,通过符号图整合语义、空间与时间关系。画面层建模角色、物体、动作及对话、旁白等文本成分,并系统连接至高层图谱,以捕捉叙事序列与抽象故事结构。在Manga109数据集的手动标注子集上,该框架支持四种代表性任务:动作检索、对话追踪、角色出现映射与时间线重建。系统不追求预测性能最优,而强调叙事建模的透明性,实现与认知理论中事件分割和视觉叙事机制一致的结构化推理。本工作推动可解释的叙事分析,为创作工具、叙事理解系统与互动媒体应用提供基础。
原文摘要 · Abstract (English)
We present a hierarchical knowledge graph framework for the structured semantic understanding of visual narratives, using comics as a representative domain for multimodal storytelling. The framework organizes narrative content across three levels-panel, event, and macro-event, by integrating symbolic graphs that encode semantic, spatial, and temporal relationships. At the panel level, it models visual elements such as characters, objects, and actions alongside textual components including dialogue and narration. These are systematically connected to higher-level graphs that capture narrative sequences and abstract story structures. Applied to a manually annotated subset of the Manga109 dataset, the framework supports interpretable symbolic reasoning across four representative tasks: action retrieval, dialogue tracing, character appearance mapping, and timeline reconstruction. Rather than prioritizing predictive performance, the system emphasizes transparency in narrative modeling and enables structured inference aligned with cognitive theories of event segmentation and visual storytelling. This work contributes to explainable narrative analysis and offers a foundation for authoring tools, narrative comprehension systems, and interactive media applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。