arXiv:2412.11668cs.CV2024-12被引 2

提出DOLPHIN模型与大規模中文手寫數據集,提升在線手寫作者檢索準確率。

Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach

  • 融合時序與頻域分析,通過門控注意力增強書寫細節表徵。
  • 在超67萬筆數據上驗證,模型顯著優於現有方法。
  • 適合手寫識別、個人身份認證等領域研究者參考。

在線手寫的普及催生了對高效手寫作者檢索系統的迫切需求,即從特定作者中精確搜索相關手寫樣本。儘管需求增長,該領域仍缺乏成熟方法和公開的大規模數據集。本文針對中文手寫短語提出新解決方案:首先設計了DOLPHIN模型,透過協同時序-頻域分析提升手寫表徵能力。頻域學習方面,提出HFGA模塊,對原始時序序列與其高頻子帶進行門控交叉注意力,強化顯著書寫細節;時序學習方面,設計CAIR模塊以促進通道交互並減少冗餘。其次,為彌補數據不足,構建了包含1,731名個人、超過67萬筆中文手寫短語的OLIWER數據集。大量實驗表明DOLPHIN在多項指標上優於現有方法。進一步探討跨領域檢索,揭示提升特徵對齊對縮小不同手寫數據分布差異至關重要。研究強調點採樣頻率與壓力特徵對表徵質量與檢索性能的關鍵影響。代碼與數據集已公開於https://github.com/SCUT-DLVCLab/DOLPHIN。

原文摘要 · Abstract (English)

Currently, the prevalence of online handwriting has spurred a critical need for effective retrieval systems to accurately search relevant handwriting instances from specific writers, known as online writer retrieval. Despite the growing demand, this field suffers from a scarcity of well-established methodologies and public large-scale datasets. This paper tackles these challenges with a focus on Chinese handwritten phrases. First, we propose DOLPHIN, a novel retrieval model designed to enhance handwriting representations through synergistic temporal-frequency analysis. For frequency feature learning, we propose the HFGA block, which performs gated cross-attention between the vanilla temporal handwriting sequence and its high-frequency sub-bands to amplify salient writing details. For temporal feature learning, we propose the CAIR block, tailored to promote channel interaction and reduce channel redundancy. Second, to address data deficit, we introduce OLIWER, a large-scale online writer retrieval dataset encompassing over 670,000 Chinese handwritten phrases from 1,731 individuals. Through extensive evaluations, we demonstrate the superior performance of DOLPHIN over existing methods. In addition, we explore cross-domain writer retrieval and reveal the pivotal role of increasing feature alignment in bridging the distributional gap between different handwriting data. Our findings emphasize the significance of point sampling frequency and pressure features in improving handwriting representation quality and retrieval performance. Code and dataset are available at https://github.com/SCUT-DLVCLab/DOLPHIN.

手寫檢索時序分析頻域建模中文文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。