用发明人原始草稿生成完整专利,解决真实专利撰写难题
Benchmarking Patent Drafting from Inventor-Style Disclosures

- 从发明人非正式草稿直接生成完整专利申请
- 新数据集Dis2Pat模拟真实专利流程,挑战长文本法律写作
- 提出可本地部署的多智能体框架Patent-MAF,性能超越开源模型
尽管近期大语言模型在单项专利撰写任务中表现良好,但它们未能解决真实专利撰写的核心挑战:直接从早期发明材料生成完整且法律一致的专利申请。以往工作多假设输入为后期、高度结构化或已法律化的文本,而实际专利流程始于发明人撰写的非正式、去法律化的披露。为弥合这一差距,我们引入Dis2Pat数据集,该数据集通过要求从发明人风格的去法律化披露生成完整的专利申请,反映了真实的专利流程。由于长文本、法律约束强且隐私要求高,我们进一步提出Patent-MAF,一个可用于本地部署的多智能体专利撰写基线框架。基准测试表明,当前LLM在专利撰写方面仍存在局限,而Patent-MAF在性能上优于评估的开源模型,并与大型闭源模型保持竞争力。
原文摘要 · Abstract (English)
While recent large language models (LLMs) have achieved promising results on individual patent drafting tasks, they fundamentally fail to investigate the core challenge of real-world patent drafting: generating a complete and legally coherent patent application directly from early-stage invention materials. Prior work predominantly assumes later-stage, highly structured, or already legalistic inputs. However, real patenting workflows begin with informal, de-legalized disclosures authored by inventors. To bridge the gap, we introduce Dis2Pat, a disclosure-to-patent dataset that reflects realistic patenting workflows by requiring the generation of complete patent applications directly from inventor-style, de-legalized disclosures. Given the inherent difficulty of long-form, legally constrained patent drafting and the strong privacy requirements, we further propose a strong baseline named Patent-MAF. It is a multi-agent framework for locally deployable patent drafting. Benchmark results reveal that current LLMs exhibit limitations in patent drafting, while Patent-MAF provides a strong baseline that consistently outperforms evaluated open-source models and remains competitive with large closed-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。