arXiv:2510.07951cs.CVcs.AI2025-10

构建超大规模动漫场景文字检测数据集,解决风格多样、布局混乱的识别难题。

A Large-scale Dataset for Robust Complex Anime Scene Text Detection

  • 聚焦动漫场景,设计包含420万标注文本块的大规模数据集
  • 在跨数据集测试中,基于该数据集训练的模型性能超越现有方法
  • 支持复杂背景与手写体,适合研究动漫文字鲁棒检测的开发者

现有文字检测数据集多针对自然或文档场景,其文字通常具有规则字体、单调色彩和有序排布,常沿直线或曲线排列。然而,动漫场景中的文字风格多样、布局不规则,且易与符号、装饰图案混淆,大量使用手写体和艺术字体。为填补这一空白,我们提出AnimeText,一个包含73.5万张图像和420万标注文本块的大规模数据集。该数据集具备层级标注和专为动漫场景设计的困难负样本。通过采用先进文字检测方法进行跨数据集评估,结果表明:在动漫场景文字检测任务中,基于AnimeText训练的模型表现优于现有数据集训练的模型。AnimeText已在HuggingFace发布:https://huggingface.co/datasets/deepghs/AnimeText

原文摘要 · Abstract (English)

Current text detection datasets primarily target natural or document scenes, where text typically appear in regular font and shapes, monotonous colors, and orderly layouts. The text usually arranged along straight or curved lines. However, these characteristics differ significantly from anime scenes, where text is often diverse in style, irregularly arranged, and easily confused with complex visual elements such as symbols and decorative patterns. Text in anime scene also includes a large number of handwritten and stylized fonts. Motivated by this gap, we introduce AnimeText, a large-scale dataset containing 735K images and 4.2M annotated text blocks. It features hierarchical annotations and hard negative samples tailored for anime scenarios. %Cross-dataset evaluations using state-of-the-art methods demonstrate that models trained on AnimeText achieve superior performance in anime text detection tasks compared to existing datasets. To evaluate the robustness of AnimeText in complex anime scenes, we conducted cross-dataset benchmarking using state-of-the-art text detection methods. Experimental results demonstrate that models trained on AnimeText outperform those trained on existing datasets in anime scene text detection tasks. AnimeText on HuggingFace: https://huggingface.co/datasets/deepghs/AnimeText

文字检测动漫图像大规模数据集视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。