构建漫画理解新基准,测试模型补全缺失画格的能力
ComicsPAP: understanding comic strips by picking the correct panel
- 设计五类任务的漫画画格补全评测框架
- 现有多模态模型表现接近随机,暴露其时序理解缺陷
- 小模型经适配后效果超越大模型,验证基准价值
大型多模态模型(LMMs)在图像描述、视觉问答和视频理解方面取得显著进展,但在漫画中复杂的时空线索面前仍显不足。为此,我们提出ComicsPAP,一个大规模漫画理解基准,包含超过10万样本,分为5个子任务,采用“选画格”范式,要求模型识别序列中缺失的画格。在多图与单图两种协议下评估显示,当前最先进的LMMs表现接近随机水平,凸显其在捕捉序列与上下文依赖方面的重大局限。为缩小差距,我们对LMMs进行适配,在ComicsPAP上取得优于10倍更大模型的效果,证明该基准能有效推动未来多模态漫画理解研究。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have made impressive strides in image captioning, VQA, and video comprehension, yet they still struggle with the intricate temporal and spatial cues found in comics. To address this gap, we introduce ComicsPAP, a large-scale benchmark designed for comic strip understanding. Comprising over 100k samples and organized into 5 subtasks under a Pick-a-Panel framework, ComicsPAP demands models to identify the missing panel in a sequence. Our evaluations, conducted under both multi-image and single-image protocols, reveal that current state-of-the-art LMMs perform near chance on these tasks, underscoring significant limitations in capturing sequential and contextual dependencies. To close the gap, we adapted LMMs for comic strip understanding, obtaining better results on ComicsPAP than 10x bigger models, demonstrating that ComicsPAP offers a robust resource to drive future research in multimodal comic comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。