arXiv:2601.04758cs.CLcs.AI2026-01中稿 · EMNLP被引 3

首个面向专利法律推理的基准,评估大模型在真实案件中的结构化判案能力。

PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks

  • 构建与PTAB裁决对齐的专利案件数据集,设计三类符合IRAC框架的分类任务。
  • 闭源模型在议题类型任务上微平均F1超0.75,开源最强模型仅达0.56,差距显著。
  • 适合关注专利AI、法律大模型评估与司法推理系统研发的研究者使用。

美国专利商标局(USPTO)的专利上诉审判委员会(PTAB)每年处理数千件单方上诉案件,需融合技术理解与法律推理。尽管大语言模型(LLMs)在专利与法律实践中应用日益广泛,但其使用仍局限于轻量级任务,缺乏系统评估其在专利领域结构化法律推理能力的方法。本文提出PILOT-Bench,首个以PTAB为中心的基准,将PTAB裁决与USPTO专利数据在案件层面进行对齐,并形式化三个符合IRAC框架的分类任务:议题类型(Issue Type)、委员会依据(Board Authorities)和子判决(Subdecision)。我们评估了多种闭源(商业)与开源LLMs,开展多角度分析,包括输入变化设置、模型族别与错误倾向。值得注意的是,在议题类型任务中,闭源模型微平均F1 consistently 超过0.75,而最强开源模型(Qwen-8B)表现约0.56,凸显推理能力的巨大差距。PILOT-Bench为专利领域法律推理的系统评估奠定基础,并指明未来通过数据设计与模型对齐提升大模型能力的方向。所有数据、代码与基准资源可在https://github.com/TeamLab/pilot-bench获取。

原文摘要 · Abstract (English)

The Patent Trial and Appeal Board (PTAB) of the USPTO adjudicates thousands of ex parte appeals each year, requiring the integration of technical understanding and legal reasoning. While large language models (LLMs) are increasingly applied in patent and legal practice, their use has remained limited to lightweight tasks, with no established means of systematically evaluating their capacity for structured legal reasoning in the patent domain. In this work, we introduce PILOT-Bench, the first PTAB-centric benchmark that aligns PTAB decisions with USPTO patent data at the case-level and formalizes three IRAC-aligned classification tasks: Issue Type, Board Authorities, and Subdecision. We evaluate a diverse set of closed-source (commercial) and open-source LLMs and conduct analyses across multiple perspectives, including input-variation settings, model families, and error tendencies. Notably, on the Issue Type task, closed-source models consistently exceed 0.75 in Micro-F1 score, whereas the strongest open-source model (Qwen-8B) achieves performance around 0.56, highlighting a substantial gap in reasoning capabilities. PILOT-Bench establishes a foundation for the systematic evaluation of patent-domain legal reasoning and points toward future directions for improving LLMs through dataset design and model alignment. All data, code, and benchmark resources are available at https://github.com/TeamLab/pilot-bench.

法律推理专利分析大模型评估IRAC框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。