解决甲骨文识别中的长尾分布问题,提升小众字识别效果。
Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and Benchmark
- 构建结构对齐的甲骨文生成数据集Oracle-P15K,含14,542张图
- 提出OBIDiff扩散模型,可精准迁移真实拓片风格到字形图
- 实验验证生成数据有效提升下游任务性能,适合古文字研究者
甲骨文识别对理解中国古代历史文化具有重要意义。然而,现有甲骨文数据集存在长尾分布问题,导致识别模型在多数类与少数类上表现不均。随着生成模型的发展,基于合成数据增强的方法成为扩充少数类样本的可行路径。但当前数据集缺乏大规模结构对齐的图像对用于生成模型训练。为此,我们首先构建了结构对齐的甲骨文生成与去噪数据集Oracle-P15K,包含14,542张融合专家领域知识的图像。其次,提出基于扩散模型的伪甲骨文生成器OBIDiff,可在给定清晰字形图和目标拓片风格图时,有效将原始拓片噪声风格迁移到字形图上。大量下游任务实验及用户偏好研究证明,Oracle-P15K与OBIDiff能准确保留字形结构并有效迁移真实拓片风格,显著提升识别性能。
原文摘要 · Abstract (English)
The oracle bone inscription (OBI) recognition plays a significant role in understanding the history and culture of ancient China. However, the existing OBI datasets suffer from a long-tail distribution problem, leading to biased performance of OBI recognition models across majority and minority classes. With recent advancements in generative models, OBI synthesis-based data augmentation has become a promising avenue to expand the sample size of minority classes. Unfortunately, current OBI datasets lack large-scale structure-aligned image pairs for generative model training. To address these problems, we first present the Oracle-P15K, a structure-aligned OBI dataset for OBI generation and denoising, consisting of 14,542 images infused with domain knowledge from OBI experts. Second, we propose a diffusion model-based pseudo OBI generator, called OBIDiff, to achieve realistic and controllable OBI generation. Given a clean glyph image and a target rubbing-style image, it can effectively transfer the noise style of the original rubbing to the glyph image. Extensive experiments on OBI downstream tasks and user preference studies show the effectiveness of the proposed Oracle-P15K dataset and demonstrate that OBIDiff can accurately preserve inherent glyph structures while transferring authentic rubbing styles effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。