用大模型自动生成专利摘要,提升撰写效率。
PATENTWRITER: A Benchmarking Study for Patent Drafting with LLMs
- 构建首个统一评测框架,测试六款大模型生成专利摘要能力。
- 大模型生成摘要质量高,部分优于专业基线,且风格符合要求。
- 适合专利从业者、AI研究者及政策制定者参考使用。
大型语言模型(LLMs)在多个重要领域展现出变革性潜力。本文旨在通过引入大模型,推动专利撰写范式革新,解决繁琐的专利申请流程问题。我们提出 PATENTWRITER,首个用于评估大模型在专利摘要生成任务中表现的统一基准框架。基于专利的第一条权利要求,我们在零样本、少样本和思维链提示策略下,对六款领先大模型(包括 GPT-4 和 LLaMA-3)进行评估。该基准不仅涵盖标准 NLP 指标(如 BLEU、ROUGE、BERTScore),还系统评估了模型在三种输入扰动下的鲁棒性,以及在专利分类与检索两项下游任务中的适用性,并开展风格分析以评估摘要长度、可读性和语气。实验表明,现代大模型能生成高保真度且风格恰当的专利摘要,常优于领域专用基线。相关代码与数据集已开源,以支持可复现性与后续研究。
原文摘要 · Abstract (English)
Large language models (LLMs) have emerged as transformative approaches in several important fields. This paper aims for a paradigm shift for patent writing by leveraging LLMs to overcome the tedious patent-filing process. In this work, we present PATENTWRITER, the first unified benchmarking framework for evaluating LLMs in patent abstract generation. Given the first claim of a patent, we evaluate six leading LLMs -- including GPT-4 and LLaMA-3 -- under a consistent setup spanning zero-shot, few-shot, and chain-of-thought prompting strategies to generate the abstract of the patent. Our benchmark PATENTWRITER goes beyond surface-level evaluation: we systematically assess the output quality using a comprehensive suite of metrics -- standard NLP measures (e.g., BLEU, ROUGE, BERTScore), robustness under three types of input perturbations, and applicability in two downstream patent classification and retrieval tasks. We also conduct stylistic analysis to assess length, readability, and tone. Experimental results show that modern LLMs can generate high-fidelity and stylistically appropriate patent abstracts, often surpassing domain-specific baselines. Our code and dataset are open-sourced to support reproducibility and future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。