arXiv:2411.00686cs.CLcs.AI2024-11NeurIPS被引 2

通过在模型层注入噪声,实现高效多样知识注入。

Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language Models

  • 在模型早期层加输入相关噪声,生成语义一致的多样化数据
  • 相比传统微调,知识注入效果提升,且无需重复生成同义句
  • 适合需要频繁更新知识的领域应用,如医疗、法律等

随着大语言模型在持续演进的知识领域中广泛应用,及时精准地注入新知识变得至关重要。尽管使用改写数据进行微调是常见方法,但仍面临两大挑战:因反复调用外部模型导致计算成本高,以及样本多样性有限。为此,我们提出LaPael,一种在潜在空间中对早期模型层施加输入依赖性噪声的改写方法。该方法可在模型内部直接生成多样且语义一致的数据增强,同时避免每次知识更新时重复生成同义句的开销。在问答基准上的大量实验表明,LaPael在知识注入效果上优于标准微调和现有基于噪声的方法。此外,将LaPael与数据层面的改写结合,可进一步提升性能。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly deployed in specialized domains with continuously evolving knowledge, the need for timely and precise knowledge injection has become essential. Fine-tuning with paraphrased data is a common approach to enhance knowledge injection, yet it faces two significant challenges: high computational costs due to repetitive external model usage and limited sample diversity. To this end, we introduce LaPael, a latent-level paraphrasing method that applies input-dependent noise to early LLM layers. This approach enables diverse and semantically consistent augmentations directly within the model. Furthermore, it eliminates the recurring costs of paraphrase generation for each knowledge update. Our extensive experiments on question-answering benchmarks demonstrate that LaPael improves knowledge injection over standard fine-tuning and existing noise-based approaches. Additionally, combining LaPael with data-level paraphrasing further enhances performance.

知识注入大模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。