arXiv:2501.11895cs.CVcs.LG2025-01

提出CMAE模型,提升未知作者手写识别准确率。

Contrastive Masked Autoencoders for Character-Level Open-Set Writer Identification

  • 结合掩码自编码与对比学习,捕捉字符级笔迹特征
  • 在CASIA数据集上达到89.7%识别精度,SOTA表现
  • 适合数字取证、文档鉴真等开放场景应用

在数字取证与文档认证领域,基于手写风格的作者识别至关重要。主要挑战是‘开放集’场景——需识别训练中未出现过的作者。代表性学习是关键,可捕捉独特笔迹特征以识别新风格。本文提出字符级开放集作者识别的对比掩码自编码器(CMAE),将掩码自编码器(MAE)与对比学习(CL)融合,同步提取序列信息并区分不同书写风格。在CASIA在线手写数据集上,模型取得89.7%的精确率,达到当前最优水平。研究推动了通用作者识别的发展,为日益互联的世界提供高效手写分析支持。

原文摘要 · Abstract (English)

In the realm of digital forensics and document authentication, writer identification plays a crucial role in determining the authors of documents based on handwriting styles. The primary challenge in writer-id is the "open-set scenario", where the goal is accurately recognizing writers unseen during the model training. To overcome this challenge, representation learning is the key. This method can capture unique handwriting features, enabling it to recognize styles not previously encountered during training. Building on this concept, this paper introduces the Contrastive Masked Auto-Encoders (CMAE) for Character-level Open-Set Writer Identification. We merge Masked Auto-Encoders (MAE) with Contrastive Learning (CL) to simultaneously and respectively capture sequential information and distinguish diverse handwriting styles. Demonstrating its effectiveness, our model achieves state-of-the-art (SOTA) results on the CASIA online handwriting dataset, reaching an impressive precision rate of 89.7%. Our study advances universal writer-id with a sophisticated representation learning approach, contributing substantially to the ever-evolving landscape of digital handwriting analysis, and catering to the demands of an increasingly interconnected world.

手写识别对比学习开放集数字取证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。