arXiv:2602.09448cs.IRcs.LG2026-02被引 1

多查询合成提升检索模型泛化能力,复杂度决定最优多样性

The Wisdom of Many Queries: Complexity-Diversity Principle for Dense Retriever Training

  • 根据查询复杂度动态调整多查询生成策略
  • 复杂度高的查询使模型在跨域任务上表现更好,提升显著
  • 适合需要强泛化能力的推理型检索场景

合成查询已成为训练密集检索器的关键,但以往方法每篇文档仅生成一个查询,仅关注查询质量。我们首次系统研究多查询合成,发现存在质量与多样性权衡:高质量查询利于领域内任务,多样查询则提升跨域(OOD)泛化能力。在Contriever、RetroMAE和Qwen3-Embedding上进行4类基准测试的控制实验表明,多样性收益与查询复杂度高度相关(r≥0.95,p<0.05),可用内容词数(CW)近似衡量。我们提出复杂度-多样性原则(CDP):查询复杂度决定最优多样性。基于此,提出复杂度感知训练:对高复杂度任务采用多查询合成,对现有数据使用CW加权训练。两种策略均提升推理密集型基准上的跨域性能,结合后效果更优。

原文摘要 · Abstract (English)

Synthetic query generation has become essential for training dense retrievers, yet prior methods generate one query per document, focusing solely on query quality. We are the first to systematically study multi-query synthesis and discover a quality-diversity trade-off: high-quality queries benefit in-domain tasks, while diverse queries benefit out-of-domain (OOD) generalization. Through controlled experiments on 4 benchmark types across Contriever, RetroMAE, and Qwen3-Embedding, we find that diversity benefit strongly correlates with query complexity (r$\geq$0.95, p<0.05), approximated by content words (CW). We formalize this as the Complexity-Diversity Principle (CDP): query complexity determines optimal diversity. Based on CDP, we propose complexity-aware training: multi-query synthesis for high-complexity tasks and CW-weighted training for existing data. Both strategies improve OOD performance on reasoning-intensive benchmarks, with compounded gains when combined.

检索增强多查询生成泛化能力复杂度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。