首个面向知识产权领域的综合性大模型评测基准,覆盖20项真实任务
IPBench: Benchmarking the Knowledge of Large Language Models in Intellectual Property
- 构建涵盖8类机制的知识产权任务体系,支持多场景评估
- 17个模型测试显示顶尖性能仅达75.8%准确率,仍有巨大提升空间
- 开源数据集与代码,助力法律与专利领域AI研究
知识产权(IP)是融合技术与法律知识的高度专业化领域,具有复杂性和强知识依赖性。近年来大语言模型在处理知识产权任务方面展现出潜力,可实现更高效的分析、理解和生成。然而现有数据集和评测基准多聚焦于专利或仅覆盖知识产权部分维度,难以反映真实应用场景。为此,我们提出IPBench,首个全面的知识产权任务分类体系与大规模双语评测基准,涵盖8类知识产权机制和20项具体任务,用于评估大模型在真实知识产权场景中的表现。我们对17个主流大模型(包括通用型与领域专用型,聊天导向与推理导向模型)在零样本、少样本及思维链设置下进行评测。结果显示,即使表现最优的DeepSeek-V3模型,准确率也仅为75.8%,表明仍有显著提升空间。值得注意的是,开源的知识产权与法律导向模型性能落后于闭源通用模型。为推动后续研究,我们公开发布IPBench,并计划持续扩展更多任务以更好反映现实复杂性,支持模型在知识产权领域的进步。相关数据与代码已通过补充链接提供。
原文摘要 · Abstract (English)
Intellectual Property (IP) is a highly specialized domain that integrates technical and legal knowledge, making it inherently complex and knowledge-intensive. Recent advancements in LLMs have demonstrated their potential to handle IP-related tasks, enabling more efficient analysis, understanding, and generation of IP-related content. However, existing datasets and benchmarks focus narrowly on patents or cover limited aspects of the IP field, lacking alignment with real-world scenarios. To bridge this gap, we introduce IPBench, the first comprehensive IP task taxonomy and a large-scale bilingual benchmark encompassing 8 IP mechanisms and 20 distinct tasks, designed to evaluate LLMs in real-world IP scenarios. We benchmark 17 main LLMs, ranging from general purpose to domain-specific, including chat-oriented and reasoning-focused models, under zero-shot, few-shot, and chain-of-thought settings. Our results show that even the top-performing model, DeepSeek-V3, achieves only 75.8% accuracy, indicating significant room for improvement. Notably, open-source IP and law-oriented models lag behind closed-source general-purpose models. To foster future research, we publicly release IPBench, and will expand it with additional tasks to better reflect real-world complexities and support model advancements in the IP domain. We provide the data and code in the supplementary URLs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。