构建大规模真实知识编辑数据集,评估大模型长期学习能力。
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
- 构建包含50万+问答对的现实世界知识编辑基准
- 实测发现现有方法难以有效处理大规模真实编辑
- 适合关注大模型持续学习与知识更新的研究者
保持大语言模型事实准确性对实际部署至关重要,但昂贵的重新训练仍是挑战。知识编辑提供了有前景的替代方案,但现有方法仅在小规模或合成数据上测试。本文提出WikiBigEdit——一个大规模真实世界维基数据编辑基准,可自动扩展以支持未来验证。首个版本包含超过50万组问答对及完整评估流程。利用该基准,我们评估现有知识编辑技术在整合大量真实事实方面的表现,并对比检索增强、持续微调等通用方法,全面揭示当前终身知识编辑的实际能力边界。
原文摘要 · Abstract (English)
Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but methods are only tested on small-scale or synthetic edit benchmarks. In this work, we aim to bridge research into lifelong knowledge editing to real-world edits at a practically relevant scale. We first introduce WikiBigEdit; a large-scale benchmark of real-world Wikidata edits, built to automatically extend lifelong for future-proof benchmarking. In its first instance, it includes over 500K question-answer pairs for knowledge editing alongside a comprehensive evaluation pipeline. Finally, we use WikiBigEdit to study existing knowledge editing techniques' ability to incorporate large volumes of real-world facts and contrast their capabilities to generic modification techniques such as retrieval augmentation and continual finetuning to acquire a complete picture of the practical extent of current lifelong knowledge editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。