arXiv:2411.17332cs.LG2024-11CVPR被引 14

研究手写文字识别模型在真实场景下的泛化能力,发现文本差异是关键瓶颈。

On the Generalization of Handwritten Text Recognition Models

  • 在无先验的域泛化设置下评估336个跨域场景
  • 发现70%情况下外分布误差可被可靠估计,偏差低于10点
  • 揭示文本差异比视觉差异更影响模型泛化,适合研究实际应用的团队

手写文字识别(HTR)近年在标准基准上取得显著进展,主要聚焦于独立同分布(i.i.d.)假设下的误差最小化。然而,这一假设在真实场景中并不成立,促使研究转向迁移学习与域适应技术。本文研究了现有HTR模型在分布外(OOD)数据上的泛化局限性,采用挑战性的域泛化设置,要求模型在无任何先验访问的情况下泛化至未知域。我们分析了来自8个先进HTR模型在7个常用数据集上的336个OOD案例,覆盖5种语言。同时考察合成数据对泛化的作用。结果表明,泛化性能主要受领域间文本差异影响,其次为视觉差异。我们证明,HTR模型在OOD场景中的误差可被可靠估计,70%情况下误差偏差低于10点。该研究揭示了当前HTR模型的根本局限,为未来研究提供了基础。

原文摘要 · Abstract (English)

Recent advances in Handwritten Text Recognition (HTR) have led to significant reductions in transcription errors on standard benchmarks under the i.i.d. assumption, thus focusing on minimizing in-distribution (ID) errors. However, this assumption does not hold in real-world applications, which has motivated HTR research to explore Transfer Learning and Domain Adaptation techniques. In this work, we investigate the unaddressed limitations of HTR models in generalizing to out-of-distribution (OOD) data. We adopt the challenging setting of Domain Generalization, where models are expected to generalize to OOD data without any prior access. To this end, we analyze 336 OOD cases from eight state-of-the-art HTR models across seven widely used datasets, spanning five languages. Additionally, we study how HTR models leverage synthetic data to generalize. We reveal that the most significant factor for generalization lies in the textual divergence between domains, followed by visual divergence. We demonstrate that the error of HTR models in OOD scenarios can be reliably estimated, with discrepancies falling below 10 points in 70\% of cases. We identify the underlying limitations of HTR models, laying the foundation for future research to address this challenge.

手写识别域泛化文本差异外分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。