构建多领域RAG评测基准并提升模型跨域泛化能力
Adapting Large Language Models for Multi-Domain Retrieval-Augmented-Generation
- 设计8个来源13个领域的多样化问答数据集
- 序列级蒸馏使跨域性能提升,优于传统微调
- 适合关注RAG系统鲁棒性的研究与工程人员
检索增强生成(RAG)可提升大语言模型的事实准确性,但多领域应用面临缺乏多样化评测基准和泛化能力差的问题。本文首个贡献是构建一个涵盖13个领域的多样化基准,包含来自8个来源的各类问答任务。第二个贡献是系统评估典型RAG微调策略在跨域场景下的表现。结果表明,标准微调难以有效泛化,而采用教师模型生成标签的序列级蒸馏方法,通过更连贯的监督信号,显著提升了跨域性能。研究揭示了增强多领域RAG鲁棒性的关键策略。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) enhances LLM factuality, but multi-domain applications face challenges like lack of diverse benchmarks and poor out-of-domain generalization. The first contribution of this work is to introduce a diverse benchmark comprising a variety of question-answering tasks from 8 sources and covering 13 domains. Our second contribution consists in systematically testing out-of-domain generalization for typical RAG tuning strategies. While our findings reveal that standard fine-tuning fails to generalize effectively, we show that sequence-level distillation with teacher-generated labels improves out-of-domain performance by providing more coherent supervision. Our findings highlight key strategies for improving multi-domain RAG robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。