arXiv:2608.16918cs.IRcs.AI2026-08

用稀疏中心向量提升专利文献检索的语义匹配效率

Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval

论文配图:Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval
图 1 · 摘自论文原文
  • 将局部文本片段映射到稀疏的语义中心,实现分段语义捕捉
  • 在CLEF-IP 2013上达到与强密集编码器相当的文档级召回率
  • 适合需要高效第一阶段检索的专利搜索场景

专利先验技术检索是一项面向长篇且高度结构化技术文档的高召回率搜索任务。稠密检索虽提升语义匹配能力,但单向量表示可能将多个技术组件、功能和约束压缩至单一嵌入。本文提出无监督的稀疏覆盖(Sparse Coverage)框架,将局部文本片段嵌入映射到嵌入空间中的稀疏词汇中心。这些中心通过面向覆盖率的k-center目标选取,片段激活邻近中心以生成兼容倒排索引检索的稀疏表示。在CLEF-IP 2013数据集上的实验表明,Sparse Coverage在多种配置下达到或超过强密集编码器的文档级召回率,同时在段落级检索中仍具竞争力。通过结合局部语义证据与稀疏倒排索引搜索,该方法为专利检索提供了高效的首阶段检索方案。

原文摘要 · Abstract (English)

Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves semantic matching, but single-vector representations may compress multiple technical components, functions, and constraints into a single embedding. We propose Sparse Coverage, an unsupervised semantic retrieval framework that maps local span embeddings to a sparse vocabulary of embedding-space centers. The centers are selected with a coverage-oriented k-center objective, and spans activate nearby centers to produce sparse representations compatible with inverted-index retrieval. Experiments on CLEF-IP 2013 show that Sparse Coverage matches or exceeds the document-level recall of strong dense patent encoders in several configurations, while remaining competitive for passage-level retrieval. By combining local semantic evidence with sparse inverted-index search, Sparse Coverage provides an effective first-stage retrieval approach for patent search.

专利检索稀疏表示语义匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。