arXiv:2505.07345cs.CLcs.AI2025-05ACL被引 1

用两个不同架构的小模型组合,让韩语搜索更准更快。

QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines

论文配图:QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines
图 1 · 摘自论文原文
  • 用生成式与嵌入式小模型融合,提升相关性判断准确率。
  • 相比顶尖大模型,推理速度提升60倍,相关性评分提高1.9%。
  • 适合需要高效率、高精度的实时搜索系统部署。

大型语言模型(LLMs)被广泛用于信息检索中的相关性评估。然而,我们的研究表明,将两种不同架构的小型语言模型(SLMs)结合使用,可在该任务上超越大模型。我们的方法QUPID融合了生成式SLM与嵌入式SLM,不仅在相关性判断准确率上表现更优,还显著降低了计算成本,相比当前最优的LLM方案更具可扩展性。在多种文档类型上的实验显示,该方法实现了更高的一致性(Cohen's Kappa 0.646,优于领先LLMs的0.387),同时推理速度提升60倍。此外,在生产搜索流水线中集成后,nDCG@5得分提升1.9%。这些结果表明,模型架构的多样性组合能显著提升信息检索系统的相关性与运行效率。

原文摘要 · Abstract (English)

Large language models (LLMs) have been widely used for relevance assessment in information retrieval. However, our study demonstrates that combining two distinct small language models (SLMs) with different architectures can outperform LLMs in this task. Our approach -- QUPID -- integrates a generative SLM with an embedding-based SLM, achieving higher relevance judgment accuracy while reducing computational costs compared to state-of-the-art LLM solutions. This computational efficiency makes QUPID highly scalable for real-world search systems processing millions of queries daily. In experiments across diverse document types, our method demonstrated consistent performance improvements (Cohen's Kappa of 0.646 versus 0.387 for leading LLMs) while offering 60x faster inference times. Furthermore, when integrated into production search pipelines, QUPID improved nDCG@5 scores by 1.9%. These findings underscore how architectural diversity in model combinations can significantly enhance both search relevance and operational efficiency in information retrieval systems.

搜索相关性小模型融合效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。