不同微调任务对大模型知识注入效果差异显著,问答类任务保留率高达48%。
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
- 以问答类任务微调可提升知识保留率至48%,远超翻译类任务的17%。
- 大模型在所有任务中均表现更好,且知识保留随模型规模呈上升趋势。
- 知识注入后在复杂场景中表现下降,说明模型未真正理解与整合信息。
随着大语言模型(LLMs)知识过时,高效更新成为迫切需求,尤其是注入专有信息时。本研究发现,需要深度理解的任务(如问答、填空)相比映射类任务(如翻译、文本转JSON),在相同事实内容下知识保留率显著更高——分别为48%与17%、20%。该现象在不同模型架构中一致存在,并遵循缩放规律:模型越大,各类任务的知识保留率越高。然而,所有模型在更广泛上下文中应用注入知识时性能均大幅下降,表明知识未能实现深层语义融合。结果强调任务选择的重要性:有效知识注入不仅依赖数据暴露,更取决于微调过程中的认知参与深度。
原文摘要 · Abstract (English)
As the knowledge of large language models (LLMs) becomes outdated over time, there is a growing need for efficient methods to update them, especially when injecting proprietary information. Our study reveals that comprehension-intensive fine-tuning tasks (e.g., question answering and blanks) achieve substantially higher knowledge retention rates (48%) compared to mapping-oriented tasks like translation (17%) or text-to-JSON conversion (20%), despite exposure to identical factual content. We demonstrate that this pattern persists across model architectures and follows scaling laws, with larger models showing improved retention across all task types. However, all models exhibit significant performance drops when applying injected knowledge in broader contexts, suggesting limited semantic integration. These findings show the importance of task selection in updating LLM knowledge, showing that effective knowledge injection relies not just on data exposure but on the depth of cognitive engagement during fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。