用多模态Transformer提升漫画分页流分割准确率
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
- 融合视觉与文本信息的多模态Transformer架构
- 在20,800页数据上实现新SOTA,F1-Macro显著提升
- 适合漫画内容分析、故事索引等下游任务研究者
本文提出CoSMo,一种用于漫画分页流分割(PSS)的新型多模态Transformer,该任务是自动化漫画内容理解的关键第一步,对角色分析、故事索引和元数据增强等下游任务至关重要。我们为这一独特媒介正式定义了PSS,并构建了一个包含20,800页的标注数据集。CoSMo采用纯视觉与多模态两种变体,在F1-Macro、全景质量及流级别指标上均优于传统基线和更大规模通用视觉-语言模型。研究发现,视觉特征在漫画分页宏观结构识别中占主导地位,但多模态信息在解决复杂歧义时仍具优势。CoSMo建立新基准,推动漫画自动化分析发展。
原文摘要 · Abstract (English)
This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a necessary first stage for many downstream tasks like character analysis, story indexing, or metadata enrichment. We formalize PSS for this unique medium and curate a new 20,800-page annotated dataset. CoSMo, developed in vision-only and multimodal variants, consistently outperforms traditional baselines and significantly larger general-purpose vision-language models across F1-Macro, Panoptic Quality, and stream-level metrics. Our findings highlight the dominance of visual features for comic PSS macro-structure, yet demonstrate multimodal benefits in resolving challenging ambiguities. CoSMo establishes a new state-of-the-art, paving the way for scalable comic book analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。