arXiv:2501.15074cs.CVcs.AI2025-01AAAI被引 9

构建专利图描述生成模型,提升专利文档理解效率。

PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures

  • 设计专用多模态视觉编码器捕捉专利图结构特征
  • 在35.5万张专利图上训练出高质量描述生成模型
  • 适合专利审查、技术翻译与知识产权保护场景

撰写专利文件中技术图纸的全面准确描述对知识共享及知识产权的复制与保护至关重要,但该任务在研究界长期被忽视。为此,我们构建了PatentDesc-355K——一个包含约35.5万张专利图及其简要和详细文本描述的大规模数据集,数据源自6万余份美国专利文档。同时,我们提出PatentLMM,一种专为专利图描述生成设计的多模态大模型。其由两个关键组件构成:(i) PatentMME,一种专门用于捕捉专利图独特结构元素的多模态视觉编码器;(ii) PatentLLaMA,基于LLaMA微调的领域适配版本,在大量专利数据上训练。实验表明,针对专利图专门设计的视觉编码器显著提升性能,生成的描述更连贯,优于同类通用多模态模型的微调效果。PatentDesc-355K与PatentLMM为自动化理解专利图铺平道路,助力高效知识共享与快速专利撰写。代码与数据已公开。

原文摘要 · Abstract (English)

Writing comprehensive and accurate descriptions of technical drawings in patent documents is crucial to effective knowledge sharing and enabling the replication and protection of intellectual property. However, automation of this task has been largely overlooked by the research community. To this end, we introduce PatentDesc-355K, a novel large-scale dataset containing ~355K patent figures along with their brief and detailed textual descriptions extracted from more than 60K US patent documents. In addition, we propose PatentLMM - a novel multimodal large language model specifically tailored to generate high-quality descriptions of patent figures. Our proposed PatentLMM comprises two key components: (i) PatentMME, a specialized multimodal vision encoder that captures the unique structural elements of patent figures, and (ii) PatentLLaMA, a domain-adapted version of LLaMA fine-tuned on a large collection of patents. Extensive experiments demonstrate that training a vision encoder specifically designed for patent figures significantly boosts the performance, generating coherent descriptions compared to fine-tuning similar-sized off-the-shelf multimodal models. PatentDesc-355K and PatentLMM pave the way for automating the understanding of patent figures, enabling efficient knowledge sharing and faster drafting of patent documents. We make the code and data publicly available.

专利分析多模态模型图像描述知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。