让大模型学新知识时仍能诚实说‘不知道’,防止胡编乱造。
SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention
- 通过稀疏微调和实体扰动正则化,防止模型在学习中遗忘‘不知道’的能力。
- 在未知问题上人类评估的拒答率提升18%至101%,同时几乎完美保留目标知识。
- 无需对齐数据或后处理,适合轻量级和隐私敏感场景使用。
大模型新增知识日益重要,但传统微调常导致其丧失‘承认未知’的能力,这在高风险场景下尤为危险。本文提出SEAT方法,通过稀疏微调控制全局激活漂移,结合实体扰动KL正则化增强局部认知边界,防止知识溢出。该方法无需对齐数据、边界探测或事后重校准,适用于轻量与隐私敏感场景。实验表明,SEAT在不同模型与数据集上使未知查询的拒答率提升18%-101%,同时保持近乎完美的目标知识掌握能力,并生成连贯、上下文感知的拒答响应。分析显示两组件均不可或缺,且能更清晰地在表示空间中分离已知与未知,同时维持下游任务性能。结果强调,保护认知拒答能力是安全知识适配的核心目标。
原文摘要 · Abstract (English)
Adapting LLMs with new knowledge is increasingly important, but standard fine-tuning often erodes aligned epistemic abstention: the ability to acknowledge when the model does not know. This failure mode is especially concerning in high-stakes settings, where abstention is a critical safeguard against hallucination. We present SEAT, a preventive fine-tuning method that preserves epistemic abstention while maintaining strong knowledge acquisition. SEAT combines sparse tuning, which constrains global activation drift, with entity-perturbed KL regularization, which sharpens local epistemic boundaries and prevents spillover to neighboring knowledge. Crucially, SEAT requires no alignment data, explicit boundary probing, or post-hoc re-alignment, making it attractive for lightweight and privacy-sensitive adaptation. Across models and datasets, SEAT improves human-evaluated abstention on unknown queries by 18%-101% over the strongest baseline while retaining near-perfect target knowledge acquisition, and produces coherent, context-aware abstentions after tuning. Further analyses show that both components are essential, that SEAT more cleanly separates known from unknown queries in representation space, and that it preserves downstream utility. These results identify preservation of epistemic abstention as a core objective for safe knowledge adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。