arXiv:2501.12751cs.IRcs.CV2025-01被引 2

用大模型提升专利图分类,支持多维度零样本学习。

Patent Figure Classification using Large Vision-language Models

  • 设计多轮选择题策略,高效处理专利图大量类别。
  • 在少样本场景下,大视觉语言模型性能优于传统CNN。
  • 构建新数据集,支持专利图类型、投影等多方面分类。

专利图分类有助于专利检索系统中的分面搜索,提升现有技术检索效率。现有方法仅针对单一维度且概念数量有限。近年来,大视觉语言模型(LVLM)在众多计算机视觉任务中表现优异,但在专利图分类领域尚未被探索。本文研究了LVLM在专利图视觉问答(VQA)和分类中的有效性,聚焦零样本与少样本学习场景。为此,我们构建了新的数据集PatFigVQA和PatFigCLS,用于多维度专利图的微调与评估(包括类型、投影、专利类别和物体)。为高效处理大量类别,提出一种基于多轮选择题的锦标赛式分类策略。在少样本设置下,基于LVLM与卷积神经网络(CNN)的多种分类方法实验表明,所提方法具有可行性。

原文摘要 · Abstract (English)

Patent figure classification facilitates faceted search in patent retrieval systems, enabling efficient prior art search. Existing approaches have explored patent figure classification for only a single aspect and for aspects with a limited number of concepts. In recent years, large vision-language models (LVLMs) have shown tremendous performance across numerous computer vision downstream tasks, however, they remain unexplored for patent figure classification. Our work explores the efficacy of LVLMs in patent figure visual question answering (VQA) and classification, focusing on zero-shot and few-shot learning scenarios. For this purpose, we introduce new datasets, PatFigVQA and PatFigCLS, for fine-tuning and evaluation regarding multiple aspects of patent figures~(i.e., type, projection, patent class, and objects). For a computational-effective handling of a large number of classes using LVLM, we propose a novel tournament-style classification strategy that leverages a series of multiple-choice questions. Experimental results and comparisons of multiple classification approaches based on LVLMs and Convolutional Neural Networks (CNNs) in few-shot settings show the feasibility of the proposed approaches.

专利分析视觉语言模型少样本学习图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。