arXiv:2507.10576cs.CYcs.AI2025-07被引 1

测试大模型在专利法实操考试中的表现,发现它们离专业律师水平还有距离。

Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?

  • 在欧洲专利律师资格考题上评估多个大模型,用真实法律任务检验能力。
  • 最佳模型GPT-o1准确率仅0.82,未达专业要求的0.90及格线。
  • 模型输出易受提示词和温度参数影响,专家更看重逻辑清晰而非答案正确。

法律领域已广泛使用大语言模型(LLMs),但其量化表现与原因仍不明确。我们评估了多个开源与专有模型——包括GPT系列、Anthropic、Deepseek和Llama-3变体——在欧洲专利律师资格考试(EQE)部分题目上的表现。OpenAI o1以0.82准确率和0.81 F1得分领先,而(亚马逊云科技)AWS Llama 3.1 8B仅得0.50准确率,部署于Python的Llama 3.1 8B为0.55,均接近二选一猜测水平。所有模型均未达到满分通过所需的0.90平均阈值,包括那些被宣传为超越博士级甚至律师水平的模型。GPT-4o在图文融合方面表现优异,而Claude 3 Opus常出现格式混乱。人类专利专家对文本论证进行评估,发现各模型存在诸多关键缺陷。他们重视表述清晰度与法律推理,而非单纯答案正确性,揭示自动指标与专家判断间的错位。模型输出对温度微调和提示词变化敏感,凸显专家监督的必要性。未来工作应聚焦逻辑一致性、鲁棒多模态与自适应提示设计,以逼近人类级专利能力。总之,尽管近期模型表现突出,公众可能高估其实际水平。构建虚拟专利律师之路依然遥远,本文旨在指出若干亟待解决的具体局限。

原文摘要 · Abstract (English)

The legal field already uses various large language models (LLMs) in actual applications, but their quantitative performance and reasons for it are underexplored. We evaluated several open-source and proprietary LLMs -- including GPT-series, Anthropic, Deepseek and Llama-3, variants -- on parts of the European Qualifying Examination (EQE) for future European Patent Attorneys. OpenAI o1 led with 0.82 accuracy and 0.81 F1 score, whereas (Amazon Web Services) AWS Llama 3.1 8B lagged at 0.50 accuracy, and a Python-deployed Llama 3.1 8B scored 0.55. The latter two are within the range of mere guessing for the two-answer forced-choice design. None of the evaluated models could have passed the examination fully, as accuracy never exceeded the average threshold of 0.90 required for professional-level standards -- also not models that are regularly promoted for their assumed beyond-PhD- and bar-admitted-lawyer-level performance. GPT-4o excelled at integrating text and graphics, while Claude 3 Opus often lost formatting coherence. Human patent experts evaluated the textual justifications and uncovered various critical shortcomings of each model. They valued clarity and legal rationale over the raw correctness of the answers, which revealed misalignment between automatic metrics and expert judgment. Model outputs were sensitive to modest temperature changes and prompt wording, which underscores the remaining necessity of expert oversight. Future work should target logical consistency, robust multimodality, and adaptive prompting to approach human-level patent proficiency. In summary, despite the outstanding performance of recent large models, the general public might overestimate their performance. The field has a long way to go to develop a virtual patent attorney. This paper wants to point out several specific limitations that need solutions.

大模型评测专利法法律AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。