arXiv:2501.07793cs.IR2025-01被引 1

无需标注数据,自动选择最佳搜索引擎提升检索生成效果

Unsupervised Query Routing for Retrieval Augmented Generation

  • 基于响应质量上界评估,无监督实现查询路由决策
  • 在5个数据集上验证,显著提升可扩展性与泛化能力
  • 适合大规模真实用户查询场景的自动化检索系统

检索增强生成中的查询路由旨在将输入查询分配给最合适的搜索引擎。现有方法严重依赖需大量人工标注的监督数据,导致成本高、可扩展性差,且对分布外场景泛化能力弱。为此,我们提出一种新型无监督方法,通过构建“上界”响应来评估检索增强响应的质量,进而决定最适合当前查询的搜索引擎。该方法无需人工标注,可自动处理大规模真实用户查询并生成训练数据。我们在五个数据集上进行了广泛实验,结果表明该方法显著提升了可扩展性与泛化能力。

原文摘要 · Abstract (English)

Query routing for retrieval-augmented generation aims to assign an input query to the most suitable search engine. Existing works rely heavily on supervised datasets that require extensive manual annotation, resulting in high costs and limited scalability, as well as poor generalization to out-of-distribution scenarios. To address these challenges, we introduce a novel unsupervised method that constructs the "upper-bound" response to evaluate the quality of retrieval-augmented responses. This evaluation enables the decision of the most suitable search engine for a given query. By eliminating manual annotations, our approach can automatically process large-scale real user queries and create training data. We conduct extensive experiments across five datasets, demonstrating that our method significantly enhances scalability and generalization capabilities.

检索增强无监督学习查询路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。