用视觉语言模型生成漫画面板密集描述,无需训练即可超越专用模型。
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
- 基于VLMs构建无训练的密集描述流水线,自动识别并关联漫画元素。
- 在13,000本漫画中标注超200万面板,生成结果在质量和数量上均领先。
- 引入保留属性的评估指标,公平评测开源VLMs,适配漫画理解任务。
漫画领域正快速发展,单页与多页分析与合成模型不断涌现。近期推出的基准和数据集支持对检测(面板、角色、文本)、关联(角色重识别与说话人识别)以及漫画元素分析(如对话转录)等任务的评估。然而,要全面理解故事情节,模型不仅需提取元素,还需理解其关系并生成高度信息丰富的描述。本文提出一个利用视觉-语言模型(VLMs)生成密集、有依据描述的流水线。为构建该流水线,我们引入一种保留属性的评估指标,用于判断描述是否涵盖所有重要属性。同时,我们创建了一个密集标注的测试集,以公平评估开源VLMs,并根据该指标选出最优描述模型。该流水线生成的带边界框的密集描述,在定量和定性上均优于专门训练的模型,且无需任何额外训练。利用该流水线,我们在13,000本书中完成了超过200万张漫画面板的标注,相关数据将发布于项目页面 https://github.com/emanuelevivoli/ComiCap。
原文摘要 · Abstract (English)
The comic domain is rapidly advancing with the development of single- and multi-page analysis and synthesis models. Recent benchmarks and datasets have been introduced to support and assess models' capabilities in tasks such as detection (panels, characters, text), linking (character re-identification and speaker identification), and analysis of comic elements (e.g., dialog transcription). However, to provide a comprehensive understanding of the storyline, a model must not only extract elements but also understand their relationships and generate highly informative captions. In this work, we propose a pipeline that leverages Vision-Language Models (VLMs) to obtain dense, grounded captions. To construct our pipeline, we introduce an attribute-retaining metric that assesses whether all important attributes are identified in the caption. Additionally, we created a densely annotated test set to fairly evaluate open-source VLMs and select the best captioning model according to our metric. Our pipeline generates dense captions with bounding boxes that are quantitatively and qualitatively superior to those produced by specifically trained models, without requiring any additional training. Using this pipeline, we annotated over 2 million panels across 13,000 books, which will be available on the project page https://github.com/emanuelevivoli/ComiCap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。