arXiv:2511.21729cs.CLcs.AI2025-11

多智能体RAG系统中,组件协同比单个组件强大更重要。

Beyond Component Strength: Synergistic Integration and Adaptive Calibration in Multi-Agent RAG Systems

  • 通过组件协同整合提升系统可靠性
  • 联合使用使拒答率从40%降至2%,无幻觉增加
  • 适合关注RAG系统部署与评估的开发者

构建可靠的检索增强生成(RAG)系统不仅需要强大的组件,更需理解组件间的交互。在50个查询(15个可回答、10个边缘案例、25个对抗性查询)上的消融实验表明,混合检索、集成验证和自适应阈值等优化单独使用时几乎无效,但联合使用可使拒答率从40%降低至2%,同时不增加幻觉。我们还发现评估难题:不同验证策略虽均安全,却可能给出不一致标签(如“拒答”与“不支持”),导致看似幻觉率异常,实为标签不一致所致。结果表明,组件协同整合比单个组件强度更重要,标准化度量与标签对正确解读性能至关重要,且即使检索质量高,也需自适应校准以避免过度自信的回答。

原文摘要 · Abstract (English)

Building reliable retrieval-augmented generation (RAG) systems requires more than adding powerful components; it requires understanding how they interact. Using ablation studies on 50 queries (15 answerable, 10 edge cases, and 25 adversarial), we show that enhancements such as hybrid retrieval, ensemble verification, and adaptive thresholding provide almost no benefit when used in isolation, yet together achieve a 95% reduction in abstention (from 40% to 2%) without increasing hallucinations. We also identify a measurement challenge: different verification strategies can behave safely but assign inconsistent labels (for example, "abstained" versus "unsupported"), creating apparent hallucination rates that are actually artifacts of labeling. Our results show that synergistic integration matters more than the strength of any single component, that standardized metrics and labels are essential for correctly interpreting performance, and that adaptive calibration is needed to prevent overconfident over-answering even when retrieval quality is high.

RAG多智能体协同优化评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。