arXiv:2505.20298cs.CLcs.AI2025-05Conference of the …被引 8

构建漫画多模态理解基准与专用模型,助力AI读懂日漫叙事

MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding

  • 设计MangaOCR和MangaVQA双基准,支持图文识别与上下文问答
  • 推出MangaLMM模型,在526组问答上实现对漫画语境的精准理解
  • 适配创作者用于故事反馈,推动多模态模型在叙事场景落地

漫画是融合图像与文本的复杂叙事形式。为提升大模型对漫画的类人理解能力,我们提出两个新基准:MangaOCR专注于页面文字识别,MangaVQA则通过526组人工构建的问答对评估上下文理解能力。基于此,我们微调开源多模态模型Qwen2.5-VL,开发出MangaLMM,可联合处理图文识别与问答任务。实验对比GPT-4o、Gemini 2.5等专有模型,验证其在漫画理解上的有效性。该基准与模型为推进多模态模型在丰富叙事领域的发展提供坚实基础。

原文摘要 · Abstract (English)

Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga creators reflect on and refine their stories. To this end, we introduce two benchmarks for multimodal manga understanding: MangaOCR, which targets in-page text recognition, and MangaVQA, a novel benchmark designed to evaluate contextual understanding through visual question answering. MangaVQA consists of 526 high-quality, manually constructed question-answer pairs, enabling reliable evaluation across diverse narrative and visual scenarios. Building on these benchmarks, we develop MangaLMM, a manga-specialized model finetuned from the open-source LMM Qwen2.5-VL to jointly handle both tasks. Through extensive experiments, including comparisons with proprietary models such as GPT-4o and Gemini 2.5, we assess how well LMMs understand manga. Our benchmark and model provide a comprehensive foundation for evaluating and advancing LMMs in the richly narrative domain of manga.

多模态漫画理解视觉问答图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。