测试大模型对植入事实的相信程度,发现只有特定方法能让模型真正‘相信’。
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
- 用泛化能力、抗质疑性、表示相似性三维度衡量事实植入深度
- 简单提示和机制编辑效果差,合成文档微调可让事实表现如真实知识
- 与常识冲突的事实仍脆弱,适合需高可信度知识的应用场景
知识编辑技术旨在向大语言模型(LLMs)植入新事实,但这些模型是否真正‘相信’这些事实?我们构建了一个衡量信念深度的框架,评估知识编辑技术的有效性。信念深度定义为:1)事实在相关上下文中的泛化能力(如经数步逻辑推演的费米估算),2)对自我审视和直接挑战的鲁棒性,3)表示方式与真实知识的相似性(通过线性探测测量)。实验表明,简单提示和机制编辑无法实现深层植入;而合成文档微调(SDF)——即在与事实一致的LLM生成文档上训练模型——常能使植入知识表现得如同真实知识。然而,若事实与基本世界知识相悖,则植入信念仍脆弱且表示上不同于真实知识。本研究提出可量化的信念深度标准,为知识编辑在真实应用中的部署提供严谨评估基础。
原文摘要 · Abstract (English)
Knowledge editing techniques promise to implant new factual knowledge into large language models (LLMs). But do LLMs really believe these facts? We develop a framework to measure belief depth and use it to evaluate the success of knowledge editing techniques. We operationalize belief depth as the extent to which implanted knowledge 1) generalizes to related contexts (e.g. Fermi estimates several logical steps removed), 2) is robust to self-scrutiny and direct challenge, and 3) is represented similarly to genuine knowledge (as measured by linear probes). Our evaluations show that simple prompting and mechanistic editing techniques fail to implant knowledge deeply. In contrast, Synthetic Document Finetuning (SDF) - where models are trained on LLM-generated documents consistent with a fact - often succeeds at implanting beliefs that behave similarly to genuine knowledge. However, SDF's success is not universal, as implanted beliefs that contradict basic world knowledge are brittle and representationally distinct from genuine knowledge. Overall, our work introduces measurable criteria for belief depth and enables the rigorous evaluation necessary for deploying knowledge editing in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。