构建首个面向专利跨领域检索的家族级基准数据集,揭示系统在跨域场景下的显著性能下降。
DAPFAM: A Domain-Aware Family-level Dataset to benchmark cross domain patent retrieval
- 按IPC3重叠定义领域边界,构建家族级跨域检索测试集
- 跨域性能比同域低约5倍,且密集检索方法难以弥补差距
- 适合关注专利信息检索鲁棒性与系统评估的研究者
当相关技术披露跨越技术边界时,专利现有技术检索变得尤为困难。现有基准缺乏明确的领域划分,难以评估检索系统应对此类变化的能力。我们提出DAPFAM,一个基于新IPC3重叠方案定义的、具有明确同域与跨域划分的家族级基准。该数据集包含1,247个查询家族和45,336个目标家族,以家族级别聚合以减少国际冗余,并采用引用作为相关性判断依据。我们进行了249组受控实验,涵盖词汇(BM25)与密集(Transformer)后端、文档与段落级检索、多种查询与文档表示、聚合策略,以及通过倒数排名融合(RRF)实现的混合融合。结果表明:跨域性能普遍比同域低约五倍,段落级检索优于文档级,密集方法虽小幅优于BM25,但无法弥合跨域差距。文档级RRF在效果与效率间取得良好平衡,开销极小。DAPFAM揭示了跨域检索的持续挑战,为开发更鲁棒的专利信息检索系统提供了可复现、计算感知的测试平台。数据集已公开于Hugging Face:https://huggingface.co/datasets/datalyes/DAPFAM_patent。
原文摘要 · Abstract (English)
Patent prior-art retrieval becomes especially challenging when relevant disclosures cross technological boundaries. Existing benchmarks lack explicit domain partitions, making it difficult to assess how retrieval systems cope with such shifts. We introduce DAPFAM, a family-level benchmark with explicit IN-domain and OUT-domain partitions defined by a new IPC3 overlap scheme. The dataset contains 1,247 query families and 45,336 target families aggregated at the family level to reduce international redundancy, with citation based relevance judgments. We conduct 249 controlled experiments spanning lexical (BM25) and dense (transformer) backends, document and passage level retrieval, multiple query and document representations, aggregation strategies, and hybrid fusion via Reciprocal Rank Fusion (RRF). Results reveal a pronounced domain gap: OUT-domain performance remains roughly five times lower than IN-domain across all configurations. Passage-level retrieval consistently outperforms document-level, and dense methods provide modest gains over BM25, but none close the OUT-domain gap. Document-level RRF yields strong effectiveness efficiency trade-offs with minimal overhead. By exposing the persistent challenge of cross-domain retrieval, DAPFAM provides a reproducible, compute-aware testbed for developing more robust patent IR systems. The dataset is publicly available on huggingface at https://huggingface.co/datasets/datalyes/DAPFAM_patent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。