arXiv:2510.15557cs.CVcs.AI2025-10中稿 · ICDAR2025 DALL

构建低资源历史影像文字识别基准,助力老旧档案文本恢复。

ClapperText: A Benchmark for Text Recognition in Low-Resource Archival Documents

  • 从二战时期拍摄视频中提取带文字的场记板,构建带精细标注的数据集。
  • 含9,813帧、9.4万词级文本,67%为手写体,1,566个部分遮挡实例。
  • 支持零样本与微调测试,适合少样本学习和历史文档分析研究者。

本文提出ClapperText,一个面向视觉退化与低资源环境下手写及印刷体文字识别的基准数据集。数据源自127段二战时期的档案视频片段,其中包含记录日期、地点、摄像师等结构化信息的场记板。数据集共包含9,813张标注帧和94,573个词级文本实例,其中67%为手写体,1,566个存在部分遮挡。每个文本实例均提供转录结果、语义类别、文本类型和遮挡状态,并以四点多边形形式标注旋转边界框,支持高精度OCR应用。识别挑战包括运动模糊、书写差异、曝光波动和背景杂乱,反映了历史文档分析中的普遍难题。提供全帧标注与裁剪词图两种格式,支持下游任务。采用统一视频级评估协议,对六种识别模型与七种检测模型在零样本与微调条件下进行评测。尽管训练集仅18段视频,微调仍带来显著性能提升,凸显该数据集在少样本学习场景中的适用性。数据集与评估代码已公开于https://github.com/linty5/ClapperText。

原文摘要 · Abstract (English)

This paper presents ClapperText, a benchmark dataset for handwritten and printed text recognition in visually degraded and low-resource settings. The dataset is derived from 127 World War II-era archival video segments containing clapperboards that record structured production metadata such as date, location, and camera-operator identity. ClapperText includes 9,813 annotated frames and 94,573 word-level text instances, 67% of which are handwritten and 1,566 are partially occluded. Each instance includes transcription, semantic category, text type, and occlusion status, with annotations available as rotated bounding boxes represented as 4-point polygons to support spatially precise OCR applications. Recognizing clapperboard text poses significant challenges, including motion blur, handwriting variation, exposure fluctuations, and cluttered backgrounds, mirroring broader challenges in historical document analysis where structured content appears in degraded, non-standard forms. We provide both full-frame annotations and cropped word images to support downstream tasks. Using a consistent per-video evaluation protocol, we benchmark six representative recognition and seven detection models under zero-shot and fine-tuned conditions. Despite the small training set (18 videos), fine-tuning leads to substantial performance gains, highlighting ClapperText's suitability for few-shot learning scenarios. The dataset offers a realistic and culturally grounded resource for advancing robust OCR and document understanding in low-resource archival contexts. The dataset and evaluation code are available at https://github.com/linty5/ClapperText.

文字识别历史档案少样本学习OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。