用图文结合方式自动生成专利文件,提升准确性和完整性。
PatentVision: A multimodal method for drafting patent applications
- 融合文本与专利图示,利用多模态模型生成完整专利说明书。
- 相比纯文本方法,输出更贴近人工撰写标准,细节更精准。
- 适合专利工程师、创新管理团队快速生成高质量专利初稿。
专利撰写因需详尽技术描述、法律合规性及视觉元素而复杂。尽管大视觉语言模型(LVLMs)在多项任务中展现潜力,其在专利自动化中的应用仍较少被探索。本文提出PatentVision,一种整合专利权利要求与图纸等多模态输入的框架,用于生成完整的专利说明书。基于先进LVLMs,PatentVision通过微调视觉语言模型并结合领域特定训练,显著提升准确性。实验表明,该方法优于纯文本方法,生成内容在保真度和与人工标准的一致性上表现更佳。引入视觉数据使其能更好呈现复杂设计特征与功能关联,产出更丰富精确的结果。本研究凸显多模态技术在专利自动化中的价值,提供可扩展工具以减轻人工负担、提升一致性。PatentVision不仅推动专利撰写革新,也为LVLM在专业领域的应用奠定基础,可能重塑知识产权管理与创新流程。
原文摘要 · Abstract (English)
Patent drafting is complex due to its need for detailed technical descriptions, legal compliance, and visual elements. Although Large Vision Language Models (LVLMs) show promise across various tasks, their application in automating patent writing remains underexplored. In this paper, we present PatentVision, a multimodal framework that integrates textual and visual inputs such as patent claims and drawings to generate complete patent specifications. Built on advanced LVLMs, PatentVision enhances accuracy by combining fine tuned vision language models with domain specific training tailored to patents. Experiments reveal it surpasses text only methods, producing outputs with greater fidelity and alignment with human written standards. Its incorporation of visual data allows it to better represent intricate design features and functional connections, leading to richer and more precise results. This study underscores the value of multimodal techniques in patent automation, providing a scalable tool to reduce manual workloads and improve consistency. PatentVision not only advances patent drafting but also lays the groundwork for broader use of LVLMs in specialized areas, potentially transforming intellectual property management and innovation processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。