针对企业搜索语义偏差,动态挖掘难负样本提升检索精度。
Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems

- 通过多模型融合与降维,动态筛选语义相近但无关的文档作为难负例。
- 在自研云服务数据集上,MRR@3提升15%,MRR@10提升19%。
- 方法可迁移至多个领域,适合知识管理与智能客服场景。
企业搜索系统常因语义不匹配和术语重叠导致无法准确获取领域特定信息,进而影响知识管理、客户支持及检索增强生成等下游应用。为此,我们提出一种专为领域内企业数据设计的可扩展难负样本挖掘框架。该方法动态选择语义挑战性强但上下文无关的文档,以增强部署的重排序模型。通过整合多种嵌入模型、执行维度压缩,并独特地选取难负例,确保计算效率与语义精确性。在自研的企业级语料库(云服务领域)上的评估显示,相比最先进基线及其他负样本采样技术,本方法在MRR@3上提升15%,在MRR@10上提升19%。在公开领域数据集(FiQA、Climate Fever、TechQA)上的进一步验证表明,该方法具备良好泛化能力,已具备真实应用条件。
原文摘要 · Abstract (English)
Enterprise search systems often struggle to retrieve accurate, domain-specific information due to semantic mismatches and overlapping terminologies. These issues can degrade the performance of downstream applications such as knowledge management, customer support, and retrieval-augmented generation agents. To address this challenge, we propose a scalable hard-negative mining framework tailored specifically for domain-specific enterprise data. Our approach dynamically selects semantically challenging but contextually irrelevant documents to enhance deployed re-ranking models. Our method integrates diverse embedding models, performs dimensionality reduction, and uniquely selects hard negatives, ensuring computational efficiency and semantic precision. Evaluation on our proprietary enterprise corpus (cloud services domain) demonstrates substantial improvements of 15\% in MRR@3 and 19\% in MRR@10 compared to state-of-the-art baselines and other negative sampling techniques. Further validation on public domain-specific datasets (FiQA, Climate Fever, TechQA) confirms our method's generalizability and readiness for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。