arXiv:2503.07519cs.IRcs.CL2025-03Conference of the …被引 3

无需分解复杂查询,70亿参数模型实现高效多跳检索

GRITHopper: Decomposition-Free Multi-Hop Dense Retrieval

  • 融合生成与表征训练,端到端优化多跳检索
  • 后检索语言建模提升信息相关性,准确率显著提高
  • 在分布内和分布外数据上均表现优异,适合实际应用

基于分解的多跳检索方法依赖大量自回归步骤拆解复杂查询,破坏端到端可微性且计算成本高。无分解方法虽缓解此问题,但在长链多跳任务和分布外数据上性能受限。为此,我们提出GRITHopper-7B,一种新型多跳密集检索模型,在分布内与分布外基准测试中均达领先水平。该模型通过融合因果语言建模与密集检索训练,结合生成与表征指令微调。受控实验表明,检索后引入额外上下文(后检索语言建模)能有效提升检索性能。通过在训练中加入最终答案等元素,模型学会更好地上下文化并检索相关信息。GRITHopper-7B提供了稳健、可扩展且泛化能力强的多跳密集检索方案,已开源供后续研究与应用使用。

原文摘要 · Abstract (English)

Decomposition-based multi-hop retrieval methods rely on many autoregressive steps to break down complex queries, which breaks end-to-end differentiability and is computationally expensive. Decomposition-free methods tackle this, but current decomposition-free approaches struggle with longer multi-hop problems and generalization to out-of-distribution data. To address these challenges, we introduce GRITHopper-7B, a novel multi-hop dense retrieval model that achieves state-of-the-art performance on both in-distribution and out-of-distribution benchmarks. GRITHopper combines generative and representational instruction tuning by integrating causal language modeling with dense retrieval training. Through controlled studies, we find that incorporating additional context after the retrieval process, referred to as post-retrieval language modeling, enhances dense retrieval performance. By including elements such as final answers during training, the model learns to better contextualize and retrieve relevant information. GRITHopper-7B offers a robust, scalable, and generalizable solution for multi-hop dense retrieval, and we release it to the community for future research and applications requiring multi-hop reasoning and retrieval capabilities.

多跳检索密集检索大模型生成式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。