研究文字数量对笔迹识别的影响,发现四行以上文本可保持90%以上准确率。
Towards the Influence of Text Quantity on Writer Retrieval
- 对比线级和字级检索,评估不同文本量下的笔迹识别性能
- 仅用一行文本时准确率下降20-30%,四行以上仍达全页90%以上
- 深度学习模型在少文本场景下显著优于传统手工特征方法
本文研究笔迹识别任务,即基于手写相似性从数据集中识别同一作者撰写的文档。现有数据集与方法多聚焦于页面级检索,本文通过评估线级和字级检索,探究文本数量对笔迹识别性能的影响。我们测试了三种前沿笔迹识别系统,涵盖手工特征与深度学习方法,并在不同文本量下进行分析。在CVL和IAM数据集上的实验表明,当仅使用一行文本作为查询和参考时,性能下降20-30%,但若至少包含四行文本,识别准确率仍能保持在全页性能的90%以上。此外,本文还证明文本依赖型检索在低文本量场景下依然表现良好。研究进一步揭示了手工特征在低文本场景中的局限性,深度学习方法如NetVLAD显著优于传统VLAD编码。
原文摘要 · Abstract (English)
This paper investigates the task of writer retrieval, which identifies documents authored by the same individual within a dataset based on handwriting similarities. While existing datasets and methodologies primarily focus on page level retrieval, we explore the impact of text quantity on writer retrieval performance by evaluating line- and word level retrieval. We examine three state-of-the-art writer retrieval systems, including both handcrafted and deep learning-based approaches, and analyze their performance using varying amounts of text. Our experiments on the CVL and IAM dataset demonstrate that while performance decreases by 20-30% when only one line of text is used as query and gallery, retrieval accuracy remains above 90% of full-page performance when at least four lines are included. We further show that text-dependent retrieval can maintain strong performance in low-text scenarios. Our findings also highlight the limitations of handcrafted features in low-text scenarios, with deep learning-based methods like NetVLAD outperforming traditional VLAD encoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。