arXiv:2410.10859cs.CLcs.AI2024-10EMNLP被引 4

构建真实多任务编辑数据集,让大模型知识修正更实用

FAME: Towards Factual Multi-Task Model Editing

  • 设计FAME数据集,覆盖真实多任务场景下的事实修正需求
  • 提出SKEME方法,用缓存机制保持模型与现实同步
  • 在多种任务中表现优异,适合实际部署中的知识更新

大型语言模型(LLMs)虽在各类任务中表现卓越,但其内部过时或错误的知识可能导致误导性输出,影响实际应用。为避免昂贵的重新训练,已有多种模型编辑方法被提出以低成本修正错误知识。然而,现有评估数据集多为单一格式的虚构数据,与真实场景脱节,实用性存疑。为此,我们提出“实用性”挑战,并构建FAME——一个真实、全面且支持多任务的事实型数据集,旨在提升模型编辑的实际适用性。进一步提出SKEME方法,采用新颖的缓存机制确保模型与现实世界同步。实验表明,SKEME在多种任务和场景下均表现优异,验证了其实际有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) embed extensive knowledge and utilize it to perform exceptionally well across various tasks. Nevertheless, outdated knowledge or factual errors within LLMs can lead to misleading or incorrect responses, causing significant issues in practical applications. To rectify the fatal flaw without the necessity for costly model retraining, various model editing approaches have been proposed to correct inaccurate knowledge within LLMs in a cost-efficient way. To evaluate these model editing methods, previous work introduced a series of datasets. However, most of the previous datasets only contain fabricated data in a single format, which diverges from real-world model editing scenarios, raising doubts about their usability in practice. To facilitate the application of model editing in real-world scenarios, we propose the challenge of practicality. To resolve such challenges and effectively enhance the capabilities of LLMs, we present FAME, an factual, comprehensive, and multi-task dataset, which is designed to enhance the practicality of model editing. We then propose SKEME, a model editing method that uses a novel caching mechanism to ensure synchronization with the real world. The experiments demonstrate that SKEME performs excellently across various tasks and scenarios, confirming its practicality.

模型编辑知识修正多任务数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。