利用专利分类层级关系提升图像检索准确率
Hierarchical Multi-Positive Contrastive Learning for Patent Image Retrieval
- 基于洛迦诺分类体系构建多正例对比损失,捕捉专利图像的层次语义
- 在DeepPatent2数据集上显著提升检索效果,尤其适合低参数模型
- 适用于资源受限环境,兼顾精度与部署效率
专利图像为表达专利创新的技术图示,专利图像检索系统需从海量数据中找出最相关图像。尽管信息检索技术不断进步,专利图像仍因技术复杂性和语义多样性带来挑战,亟需高效领域适配方法。现有方法忽略专利的层级结构,如洛迦诺国际分类(LIC)系统将大类(如“家具”)细分为子类(如“座椅”、“床”),并进一步划分至具体设计。本文提出一种分层多正例对比学习方法,利用LIC分类体系在检索过程中引入层级关联。该方法为每张图像在批次内分配多个正例对,相似度按层级关系动态调整。在DeepPatent2数据集上,结合多种视觉与多模态模型的实验表明,该方法有效提升检索性能。值得注意的是,该方法在低参数模型上表现优异,计算开销小,适合在硬件受限环境下部署。
原文摘要 · Abstract (English)
Patent images are technical drawings that convey information about a patent's innovation. Patent image retrieval systems aim to search in vast collections and retrieve the most relevant images. Despite recent advances in information retrieval, patent images still pose significant challenges due to their technical intricacies and complex semantic information, requiring efficient fine-tuning for domain adaptation. Current methods neglect patents' hierarchical relationships, such as those defined by the Locarno International Classification (LIC) system, which groups broad categories (e.g., "furnishing") into subclasses (e.g., "seats" and "beds") and further into specific patent designs. In this work, we introduce a hierarchical multi-positive contrastive loss that leverages the LIC's taxonomy to induce such relations in the retrieval process. Our approach assigns multiple positive pairs to each patent image within a batch, with varying similarity scores based on the hierarchical taxonomy. Our experimental analysis with various vision and multimodal models on the DeepPatent2 dataset shows that the proposed method enhances the retrieval results. Notably, our method is effective with low-parameter models, which require fewer computational resources and can be deployed on environments with limited hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。