arXiv:2411.13477cs.CLcs.AI2024-11被引 6

用文本蕴含判断专利新颖性,让大模型预测修改方案。

PatentEdits: Framing Patent Novelty as Textual Entailment

  • 将专利修改任务转化为文本蕴含判断,逐句分析引用文献与原文关系。
  • 在10.5万条专利修订数据上验证,大模型可有效预测哪些权利要求需修改。
  • 适合专利审查、AI辅助法律写作的研究者和从业者使用。

专利要获得美国专利商标局(USPTO)授权,必须具备新颖性和非显而易见性。若不满足,审查员会引用现有技术(prior art)驳回申请,并发出非最终驳回通知。预测哪些权利要求需修改以克服新颖性质疑,是保障发明权的关键步骤,但此前未被当作可学习的任务研究。本文构建了PatentEdits数据集,包含10.5万条成功修订案例。我们设计算法对每句话的修改进行标注,并评估大语言模型(LLMs)在预测这些修改上的表现。结果表明,通过比较引用文献与原始句子之间的文本蕴含关系,能有效预测哪些发明权利要求保持不变或相对于现有技术具有新颖性。

原文摘要 · Abstract (English)

A patent must be deemed novel and non-obvious in order to be granted by the US Patent Office (USPTO). If it is not, a US patent examiner will cite the prior work, or prior art, that invalidates the novelty and issue a non-final rejection. Predicting what claims of the invention should change given the prior art is an essential and crucial step in securing invention rights, yet has not been studied before as a learnable task. In this work we introduce the PatentEdits dataset, which contains 105K examples of successful revisions that overcome objections to novelty. We design algorithms to label edits sentence by sentence, then establish how well these edits can be predicted with large language models (LLMs). We demonstrate that evaluating textual entailment between cited references and draft sentences is especially effective in predicting which inventive claims remained unchanged or are novel in relation to prior art.

专利生成文本蕴含大模型应用法律AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。