测试RAG系统在教材知识更新下的稳定性,发现其性能显著下降。
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
- 构建模拟教材知识更新的问答数据集KnowShiftQA
- 3005个问题显示多数RAG系统性能大幅下降
- 融合课本与模型知识的问题尤其难应对
检索增强生成(RAG)系统在中小学教育领域问答中展现巨大潜力,因其通常依赖权威教材中的限定知识。然而,教材知识与大型语言模型(LLM)固有参数知识之间的差异会削弱RAG系统的有效性。为系统研究此类知识差异对RAG鲁棒性的影响,我们提出KnowShiftQA。该新型问答数据集通过人为假设性知识更新,对答案和源文档进行修改,模拟教材知识的演变。KnowShiftQA包含5个学科共3,005个问题,采用全面的问题类型体系,聚焦上下文利用与知识整合能力。大量实验表明,当面临知识差异时,大多数RAG系统性能显著下降。尤其对于需要结合上下文(教材)知识与参数化(LLM)知识的问题,当前大模型仍面临严峻挑战。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems show remarkable potential as question answering tools in the K-12 Education domain, where knowledge is typically queried within the restricted scope of authoritative textbooks. However, discrepancies between these textbooks and the parametric knowledge inherent in Large Language Models (LLMs) can undermine the effectiveness of RAG systems. To systematically investigate RAG system robustness against such knowledge discrepancies, we introduce KnowShiftQA. This novel question answering dataset simulates these discrepancies by applying deliberate hypothetical knowledge updates to both answers and source documents, reflecting how textbook knowledge can shift. KnowShiftQA comprises 3,005 questions across five subjects, designed with a comprehensive question typology focusing on context utilization and knowledge integration. Our extensive experiments on retrieval and question answering performance reveal that most RAG systems suffer a substantial performance drop when faced with these knowledge discrepancies. Furthermore, questions requiring the integration of contextual (textbook) knowledge with parametric (LLM) knowledge pose a significant challenge to current LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。