通过语义归一化构建层次化叙事知识图谱,提升漫画理解的连贯性与鲁棒性。
Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
- 基于词汇相似性和嵌入聚类合并语义相近的事件和动作
- 在Manga109数据集上提升动作检索、角色定位等任务性能
- 适合关注多模态叙事理解与知识图谱去噪的研究者
理解漫画等视觉叙事需要结构化表示来捕捉事件、角色及其跨多层级故事组织的关系。然而,符号化叙事图常存在不一致与冗余问题,相同动作或事件在不同标注或上下文中被不同命名,这种差异限制了推理与泛化能力。本文提出一种面向分层叙事知识图谱的语义归一化框架。基于认知驱动的叙事理解模型,采用词汇相似性与基于嵌入的聚类方法合并语义相关的行为与事件。该归一化过程减少标注噪声,对齐各叙事层级的符号类别,同时保持可解释性。我们在Manga109数据集的标注漫画故事上验证框架,应用于面板级、事件级与故事级图谱。初步评估显示,在动作检索、角色定位与事件摘要等叙事推理任务中,语义归一化显著提升连贯性与鲁棒性,同时维持符号透明性。结果表明,归一化是构建可扩展、认知启发式多模态叙事理解图谱的关键步骤。
原文摘要 · Abstract (English)
Understanding visual narratives such as comics requires structured representations that capture events, characters, and their relations across multiple levels of story organization. However, symbolic narrative graphs often suffer from inconsistency and redundancy, where similar actions or events are labeled differently across annotations or contexts. Such variance limits the effectiveness of reasoning and generalization. This paper introduces a semantic normalization framework for hierarchical narrative knowledge graphs. Building on cognitively grounded models of narrative comprehension, we propose methods that consolidate semantically related actions and events using lexical similarity and embedding-based clustering. The normalization process reduces annotation noise, aligns symbolic categories across narrative levels, and preserves interpretability. We demonstrate the framework on annotated manga stories from the Manga109 dataset, applying normalization to panel-, event-, and story-level graphs. Preliminary evaluations across narrative reasoning tasks, such as action retrieval, character grounding, and event summarization, show that semantic normalization improves coherence and robustness, while maintaining symbolic transparency. These findings suggest that normalization is a key step toward scalable, cognitively inspired graph models for multimodal narrative understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。