评测视觉语言模型看漫画讲故事的能力,发现它们缺乏连贯叙事理解。
Re:Verse -- Can Your VLM Read a Manga?
- 构建多模态标注框架,关联漫画画面与小说文本
- 在11章《Re:Zero》中测试308个分镜,发现模型难懂因果关系
- 提出新评估方法,适合研究长序列故事理解的学者
当前视觉语言模型在处理连续视觉叙事时,存在表面识别与深层叙事推理之间的显著差距。通过对漫画叙事理解的全面研究,我们发现尽管近期大型多模态模型在单个分镜解读上表现优异,但在时间因果和跨分镜连贯性方面系统性失败,而这些正是完整故事理解的核心要求。我们提出一种新型评估框架,结合细粒度多模态标注、跨模态嵌入分析和检索增强评估,系统刻画这些局限性。该方法包括:(i) 通过匹配轻小说文本建立视觉元素与叙事结构的严谨标注协议;(ii) 在多种推理范式下进行综合评估,涵盖直接推理与检索增强生成;(iii) 跨模态相似性分析揭示当前模型联合表示中的根本错位。将该框架应用于《Re:Zero》漫画的11章共308个分镜,首次系统研究了多模态模型在长篇叙事理解方面的表现,从生成叙事、语境对话定位和时间推理三个核心维度展开。结果表明,当前模型缺乏真正的故事级智能,尤其在非线性叙事、角色一致性及长序列因果推理方面表现不佳。本工作为评估叙事智能奠定了基础并提供了可操作的方法,为深入理解离散视觉叙事中超越基本识别的深度序列理解能力提供了重要洞察。
原文摘要 · Abstract (English)
Current Vision Language Models (VLMs) demonstrate a critical gap between surface-level recognition and deep narrative reasoning when processing sequential visual storytelling. Through a comprehensive investigation of manga narrative understanding, we reveal that while recent large multimodal models excel at individual panel interpretation, they systematically fail at temporal causality and cross-panel cohesion, core requirements for coherent story comprehension. We introduce a novel evaluation framework that combines fine-grained multimodal annotation, cross-modal embedding analysis, and retrieval-augmented assessment to systematically characterize these limitations. Our methodology includes (i) a rigorous annotation protocol linking visual elements to narrative structure through aligned light novel text, (ii) comprehensive evaluation across multiple reasoning paradigms, including direct inference and retrieval-augmented generation, and (iii) cross-modal similarity analysis revealing fundamental misalignments in current VLMs' joint representations. Applying this framework to Re:Zero manga across 11 chapters with 308 annotated panels, we conduct the first systematic study of long-form narrative understanding in VLMs through three core evaluation axes: generative storytelling, contextual dialogue grounding, and temporal reasoning. Our findings demonstrate that current models lack genuine story-level intelligence, struggling particularly with non-linear narratives, character consistency, and causal inference across extended sequences. This work establishes both the foundation and practical methodology for evaluating narrative intelligence, while providing actionable insights into the capability of deep sequential understanding of Discrete Visual Narratives beyond basic recognition in Multimodal Models. Project Page: https://re-verse.vercel.app
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。