arXiv:2604.10745cs.CL2026-04被引 2

提出首个查询变体基准,揭示自适应检索增强生成对表面变化敏感

How You Ask Matters! Adaptive RAG Robustness to Query Variations

论文配图:How You Ask Matters! Adaptive RAG Robustness to Query Variations
图 1 · 摘自论文原文
  • 构建大规模语义相同但表达多样的查询变体数据集
  • 发现微小查询变化导致检索行为与准确率大幅波动
  • 适合关注模型鲁棒性与实际应用可靠性的研究者

自适应检索增强生成(Adaptive RAG)通过动态触发检索提升效率与准确性,广泛应用于实际场景。然而,真实查询在表层形式上存在多样变化,而其对自适应RAG的影响尚未被充分研究。本文首次构建了大规模的多样化但语义一致的查询变体基准,融合人工撰写与模型生成的改写版本。该基准支持对自适应RAG关键组件的系统评估,涵盖答案质量、计算成本与检索决策三个维度。实验发现显著的鲁棒性差距:查询表面形式的微小变化会引发检索行为和准确率的剧烈波动。尽管大模型表现更优,但其鲁棒性并未随之提升。结果表明,自适应RAG方法对保持语义一致的查询变化极为脆弱,暴露出关键的鲁棒性挑战。

原文摘要 · Abstract (English)

Adaptive Retrieval-Augmented Generation (RAG) promises accuracy and efficiency by dynamically triggering retrieval only when needed and is widely used in practice. However, real-world queries vary in surface form even with the same intent, and their impact on Adaptive RAG remains under-explored. We introduce the first large-scale benchmark of diverse yet semantically identical query variations, combining human-written and model-generated rewrites. Our benchmark facilitates a systematic evaluation of Adaptive RAG robustness by examining its key components across three dimensions: answer quality, computational cost, and retrieval decisions. We discover a critical robustness gap, where small surface-level changes in queries dramatically alter retrieval behavior and accuracy. Although larger models show better performance, robustness does not improve accordingly. These findings reveal that Adaptive RAG methods are highly vulnerable to query variations that preserve identical semantics, exposing a critical robustness challenge.

自适应RAG查询鲁棒性检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。