通过知识注入与分层推理提升搜索相关性判断精度
RAG-Match: Retrieval-Augmented Knowledge Injection and Hierarchical Reasoning for Calibrated Semantic Relevance

- 三阶段框架:先强化查询语义定位,再对齐结构化推理,最后校准决策边界
- 在真实搜索数据集上优于主流大模型基线,多指标显著提升
- 适合需要精准理解用户意图和事实等价性的搜索系统研发者
在知识密集型搜索场景中,语义相关性判断尤为困难,准确排序不仅需语义匹配,还需背景支撑、多步推理及可校准的决策边界。现有相关性模型多依赖直接标签监督或浅层语义相似度,难以处理隐含意图、事实等价性及细粒度相关性差异。为此,我们提出RAG-Match,一个三阶段框架,融合知识增强预训练、分层推理对齐与基于偏好的决策校准,实现相关性建模。核心思路是先强化以查询为中心的语义定位,再对齐结构化相关性推理路径,最后修正复杂边界情形下的决策不一致问题。在真实世界搜索相关性基准上的实验表明,RAG-Match在多个排序指标上持续优于强大多模型基线,验证了知识注入、推理监督与偏好优化结合在细粒度相关性判断中的有效性。
原文摘要 · Abstract (English)
Semantic relevance judgment for search is particularly challenging in knowledge-intensive scenarios, where accurate ranking requires not only semantic matching but also background grounding, multi-step reasoning, and well-calibrated decision boundaries. Existing relevance models mainly rely on direct label supervision or shallow semantic similarity, which limits their ability to handle implicit intent, factual equivalence, and fine-grained relevance distinctions. To address this issue, we propose \textsc{RAG-Match}, a three-stage framework that integrates knowledge-augmented pretraining, hierarchical reasoning alignment, and preference-based decision calibration for relevance modeling. The key idea is to first strengthen query-centered semantic grounding, then align the model with structured relevance reasoning, and finally correct decision-level inconsistencies in difficult boundary cases. Experimental results on a real-world search relevance benchmark show that \textsc{RAG-Match} consistently outperforms strong LLM-based baselines across multiple ranking metrics, demonstrating the effectiveness of combining knowledge injection, reasoning supervision, and preference optimization for fine-grained relevance judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。