修复漫画数据集标注错误,提升AI理解漫画的准确性
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding

- 通过OCR检测与人工修正结合,重校2.9万条对话标注
- 解决转录不准、文本遗漏、气泡分割过粗等问题
- 适合研究漫画理解、OCR及多模态AI的学者使用
漫画是一种具有文化特色的多模态媒介,也是日本流行文化中最具影响力的形式之一。随着AI系统日益关注漫画理解与翻译任务,Manga109已成为相关研究的基础数据集。然而,现有Manga109数据集存在转录不准确和标注粗略的问题,难以匹配现代OCR与多模态漫画理解需求。本文重新审视Manga109的对话文本标注,识别出五类标注问题,包括转录错误、文本区域遗漏、对话与拟声词重叠、气泡分割不足等。为此,我们结合基于OCR的问题检测与人工修订,构建了Manga109-v2026,共修订约29,000条对话标注。新版本更契合现代OCR与多模态漫画理解系统,同时保留了漫画特有的表达结构。
原文摘要 · Abstract (English)
Manga is a culturally distinctive multimodal medium and one of the most influential forms of Japanese popular culture. As AI systems increasingly target manga understanding, OCR, and translation, Manga109 has become a foundational dataset for manga-related AI research. However, the current Manga109 dataset contains inaccurate transcriptions and coarse annotations, which do not align well with modern OCR and multimodal manga understanding tasks. In this work, we revisit the dialogue text annotations of Manga109 and identify five categories of annotation issues, including inaccurate transcriptions, missing text regions, overlapping dialogue and onomatopoeia, and under-segmented speech balloons. To address these issues, we combine OCR-based issue detection and manual revision to construct Manga109-v2026, revising approximately 29,000 dialogue annotations. Our revisions better align Manga109 with modern OCR and multimodal manga understanding systems while preserving expressive structures characteristic of manga.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。