arXiv:2508.15297cs.CVcs.AI2025-08EMNLP被引 1

用CLIP模型提升设计专利的图像理解与检索能力

DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding

  • 基于CLIP构建多模态框架,融合图文对比学习
  • 在专利分类与检索任务中均超越现有模型表现
  • 适合专利分析、创新设计辅助等场景使用

在设计专利分析领域,传统任务如专利分类和图像检索高度依赖图像数据。然而,专利图像通常为包含发明抽象结构元素的草图,难以充分传递视觉上下文与语义信息,可能导致现有技术检索时产生歧义。近年来,视觉语言模型(如CLIP)为更可靠、准确的AI驱动专利分析提供了可能。本文提出DesignCLIP框架,利用大规模美国设计专利数据集,结合类别感知分类与对比学习,通过生成的详细图像描述和多视图图像学习策略,提升对专利数据的理解。实验验证了DesignCLIP在专利分类、专利检索及跨模态检索等下游任务中的有效性。结果表明,该模型在所有任务中均持续优于基线与当前最优模型。研究证实了多模态方法在推动专利分析发展方面的潜力。代码库已公开:https://anonymous.4open.science/r/PATENTCLIP-4661/README.md。

原文摘要 · Abstract (English)

In the field of design patent analysis, traditional tasks such as patent classification and patent image retrieval heavily depend on the image data. However, patent images -- typically consisting of sketches with abstract and structural elements of an invention -- often fall short in conveying comprehensive visual context and semantic information. This inadequacy can lead to ambiguities in evaluation during prior art searches. Recent advancements in vision-language models, such as CLIP, offer promising opportunities for more reliable and accurate AI-driven patent analysis. In this work, we leverage CLIP models to develop a unified framework DesignCLIP for design patent applications with a large-scale dataset of U.S. design patents. To address the unique characteristics of patent data, DesignCLIP incorporates class-aware classification and contrastive learning, utilizing generated detailed captions for patent images and multi-views image learning. We validate the effectiveness of DesignCLIP across various downstream tasks, including patent classification and patent retrieval. Additionally, we explore multimodal patent retrieval, which provides the potential to enhance creativity and innovation in design by offering more diverse sources of inspiration. Our experiments show that DesignCLIP consistently outperforms baseline and SOTA models in the patent domain on all tasks. Our findings underscore the promise of multimodal approaches in advancing patent analysis. The codebase is available here: https://anonymous.4open.science/r/PATENTCLIP-4661/README.md.

多模态专利分析CLIP图像检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。