将漫画自动转为有文学感的文本,让视障者也能体验漫画故事魅力。
From Panels to Prose: Generating Literary Narratives from Comics
- 用统一模型Magiv3识别漫画中的画面、角色、文字和对话气泡。
- 构建3300+张日漫面板的标注数据集,评测视觉语言模型理解能力。
- 结合大模型生成流畅文学叙事,提升视障读者的阅读体验。
漫画作为广受欢迎的叙事形式,凭借其视觉吸引力吸引全球观众,但其强视觉特性对视障读者构成显著障碍。本文提出一种自动化系统,将日式漫画转化为富有表现力的文本叙事,使视障用户能感知角色深度、互动关系与生动场景。主要贡献包括:(1)提出统一模型Magiv3,擅长定位漫画面板、角色、文本及对话气泡尾部,执行OCR与角色定位;(2)发布超过3300个日本漫画面板的人工标注图像描述与角色定位数据,用于评估大视觉语言模型在漫画理解上的表现;(3)展示如何将大视觉语言模型与Magiv3结合,生成连贯且富有文学性的叙述文本,实现视障人群对漫画叙事丰富性的可及性。
原文摘要 · Abstract (English)
Comics have long been a popular form of storytelling, offering visually engaging narratives that captivate audiences worldwide. However, the visual nature of comics presents a significant barrier for visually impaired readers, limiting their access to these engaging stories. In this work, we provide a pragmatic solution to this accessibility challenge by developing an automated system that generates text-based literary narratives from manga comics. Our approach aims to create an evocative and immersive prose that not only conveys the original narrative but also captures the depth and complexity of characters, their interactions, and the vivid settings in which they reside. To this end we make the following contributions: (1) We present a unified model, Magiv3, that excels at various functional tasks pertaining to comic understanding, such as localising panels, characters, texts, and speech-bubble tails, performing OCR, grounding characters etc. (2) We release human-annotated captions for over 3300 Japanese comic panels, along with character grounding annotations, and benchmark large vision-language models in their ability to understand comic images. (3) Finally, we demonstrate how integrating large vision-language models with Magiv3, can generate seamless literary narratives that allows visually impaired audiences to engage with the depth and richness of comic storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。