arXiv:2510.22264cs.CLcs.AI2025-10被引 4

构建专利文本嵌入的综合性评测基准与模型家族,提升专利检索与分析能力。

PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding

  • 设计15项任务覆盖专利检索、分类等场景,含206万样本和领域分层划分
  • patembed系列模型在大模型参数量达3.44亿,支持4096词长上下文
  • 多任务训练提升泛化性,适合需要高精度专利分析的研究与工业应用

专利文本嵌入可支持现有技术检索、技术布局分析与专利评估,但现有评测基准未能充分反映专利特有的挑战。我们提出PatenTEB,一个包含15项任务的综合性基准,涵盖检索、分类、重述与聚类,共206万条数据。该基准采用领域分层划分、领域特定难负样本挖掘,并系统覆盖通用嵌入基准缺失的非对称片段-文档匹配场景。通过多任务训练,我们构建了patembed模型家族,参数量从6700万至3.44亿,支持最长4096个词的上下文。外部验证显示:patembed-base在MTEB BigPatentClustering.v2上取得0.494的V-measure(前最佳0.445),patembed-large在DAPFAM上达到0.377 NDCG@100。系统消融实验表明,多任务训练虽小幅增加基准开销,但显著提升外部泛化能力;领域预训练初始化在各类任务中均带来稳定优势。所有资源将开源至https://github.com/iliass-y/patenteb。

原文摘要 · Abstract (English)

Patent text embeddings enable prior art search, technology landscaping, and patent analysis, yet existing benchmarks inadequately capture patent-specific challenges. We introduce PatenTEB, a comprehensive benchmark comprising 15 tasks across retrieval, classification, paraphrase, and clustering, with 2.06 million examples. PatenTEB employs domain-stratified splits, domain specific hard negative mining, and systematic coverage of asymmetric fragment-to-document matching scenarios absent from general embedding benchmarks. We develop the patembed model family through multi-task training, spanning 67M to 344M parameters with context lengths up to 4096 tokens. External validation shows strong generalization: patembed-base achieves state-of-the-art on MTEB BigPatentClustering.v2 (0.494 V-measure vs. 0.445 previous best), while patembed-large achieves 0.377 NDCG@100 on DAPFAM. Systematic ablations reveal that multi-task training improves external generalization despite minor benchmark costs, and that domain-pretrained initialization provides consistent advantages across task families. All resources will be made available at https://github.com/iliass-y/patenteb. Keywords: patent retrieval, sentence embeddings, multi-task learning, asymmetric retrieval, benchmark evaluation, contrastive learning.

专利分析文本嵌入多任务学习评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。