为手写生成设计更精准的评估指标,提升真实可用性判断。
Rethinking HTG Evaluation: Bridging Generation and Recognition
- 用识别模型误差设计三项专用评估指标,聚焦风格、内容与多样性。
- 在IAM数据集上验证,传统指标如FID无法准确衡量生成样本多样性与实用性。
- 适合关注手写生成质量评估、手写识别模型优化的研究者使用。
自然图像生成模型的评估已得到广泛研究,但手写生成(HTG)任务虽具特殊性,仍沿用通用评估协议与指标,可能并不恰当。本文提出三项专用于HTG评估的新指标:$\text{HTG}_{\text{HTR}}$、$\text{HTG}_{\text{style}}$ 和 $\text{HTG}_{\text{OOV}}$,其基于手写文本识别(HTR)与作者识别模型的识别准确率,强调书写风格、文本内容与多样性三个核心维度。我们在IAM手写数据库上开展全面实验,发现广泛使用的FID等指标无法有效量化生成样本的多样性与实际应用价值。结果表明,新指标信息量更丰富,凸显了建立标准化评估协议的必要性。所提评估方法更具鲁棒性与信息性,有助于提升HTR性能。代码已开源:https://github.com/koninik/HTG_evaluation。
原文摘要 · Abstract (English)
The evaluation of generative models for natural image tasks has been extensively studied. Similar protocols and metrics are used in cases with unique particularities, such as Handwriting Generation, even if they might not be completely appropriate. In this work, we introduce three measures tailored for HTG evaluation, $ \text{HTG}_{\text{HTR}} $, $ \text{HTG}_{\text{style}} $, and $ \text{HTG}_{\text{OOV}} $, and argue that they are more expedient to evaluate the quality of generated handwritten images. The metrics rely on the recognition error/accuracy of Handwriting Text Recognition and Writer Identification models and emphasize writing style, textual content, and diversity as the main aspects that adhere to the content of handwritten images. We conduct comprehensive experiments on the IAM handwriting database, showcasing that widely used metrics such as FID fail to properly quantify the diversity and the practical utility of generated handwriting samples. Our findings show that our metrics are richer in information and underscore the necessity of standardized evaluation protocols in HTG. The proposed metrics provide a more robust and informative protocol for assessing HTG quality, contributing to improved performance in HTR. Code for the evaluation protocol is available at: https://github.com/koninik/HTG_evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。