首份漫画理解任务框架,梳理数据集与核心挑战
One missing piece in Vision and Language: A Survey on Comics Understanding
- 提出漫画理解分层框架LoCU,重构视觉语言任务
- 系统整理现有漫画数据集与任务类型,覆盖多模态理解
- 适合对跨模态推理、叙事理解感兴趣的科研者
视觉语言模型已发展为在文档理解、视觉问答和定位等任务中表现优异的通用系统,常在零样本设置下运行。漫画理解作为复杂且多面的领域,有望从这些进展中获益。漫画融合丰富视觉与文本叙事,挑战模型完成图像分类、目标检测、实例分割及基于序列面板的深层叙事理解。然而,漫画独特的结构——包括风格变化、阅读顺序多样性和非线性叙事——带来了与其他视觉语言领域不同的挑战。本文综述了漫画理解的多维度研究,贡献包括:(1) 分析漫画媒介结构,详述其独特构成元素;(2) 梳理常用数据集与任务,强调其推动作用;(3) 提出漫画理解层次框架(LoCU),重新定义漫画中的视觉语言任务;(4) 基于该框架详细回顾并分类现有方法;(5) 指出现有研究挑战,并提出未来方向,尤其关注视觉语言模型在漫画中的应用。本综述是首个提出面向任务的漫画智能框架的研究,旨在通过填补数据与任务定义的空白,指导未来研究。相关项目见:https://github.com/emanuelevivoli/awesome-comics-understanding。
原文摘要 · Abstract (English)
Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering, and grounding, often in zero-shot settings. Comics Understanding, a complex and multifaceted field, stands to greatly benefit from these advances. Comics, as a medium, combine rich visual and textual narratives, challenging AI models with tasks that span image classification, object detection, instance segmentation, and deeper narrative comprehension through sequential panels. However, the unique structure of comics -- characterized by creative variations in style, reading order, and non-linear storytelling -- presents a set of challenges distinct from those in other visual-language domains. In this survey, we present a comprehensive review of Comics Understanding from both dataset and task perspectives. Our contributions are fivefold: (1) We analyze the structure of the comics medium, detailing its distinctive compositional elements; (2) We survey the widely used datasets and tasks in comics research, emphasizing their role in advancing the field; (3) We introduce the Layer of Comics Understanding (LoCU) framework, a novel taxonomy that redefines vision-language tasks within comics and lays the foundation for future work; (4) We provide a detailed review and categorization of existing methods following the LoCU framework; (5) Finally, we highlight current research challenges and propose directions for future exploration, particularly in the context of vision-language models applied to comics. This survey is the first to propose a task-oriented framework for comics intelligence and aims to guide future research by addressing critical gaps in data availability and task definition. A project associated with this survey is available at https://github.com/emanuelevivoli/awesome-comics-understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。