大模型判断对错的能力依赖表面相似性,一有变化就失效。
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- 用语义不变的扰动测试模型对真假陈述的内部表征
- 当输入偏离训练数据时,真假区分能力下降超60%
- 适合关注模型可靠性与对抗鲁棒性的研究者
为使大语言模型(LLMs)可靠,其知识表征需具备泛化能力,能应对训练中未见的多样场景。然而,已有研究表明,模型性能易受微小输入变化影响,表现出脆弱性。本文探究这种脆弱性是否源于内部知识表征的不稳定性。基于先前工作表明LLM表征可编码陈述真伪性——即真实与虚假陈述在表征空间中可被轻易分离——我们通过施加语义保持但形式改变的扰动(如拼写错误、改写),评估表征分离能力在分布外(OOD)样本上的退化情况。实验覆盖四类主流LLM、五个评测数据集和三种知识探针方法。结果表明,当陈述呈现形式与预训练数据差异增大时,真假表征的可分性显著下降;尽管在接近训练数据的形式下仍能准确区分真假,但该能力严重依赖于表面形式。这揭示了模型表现脆性的可能成因:学习到的是浅层、非鲁棒的知识表征,泛化能力有限。本研究挑战了现有真伪探针的有效性,呼吁进一步提升知识表征的鲁棒性。
原文摘要 · Abstract (English)
For Large Language Models (LLMs) to be reliable, they must learn robust knowledge that can be generally applied in diverse settings -- often unlike those seen during training. Yet, extensive research has shown that LLM performance can be brittle, with models exhibiting excessive sensitivity to trivial input variations. In this work, we explore whether this brittleness is a direct result of unstable internal knowledge representations. To explore this question, we build on previous work showing that LLM representations encode statement truthfulness -- i.e., true, factual statements can be easily separated from false, inaccurate ones. Specifically, we test the robustness of learned knowledge by evaluating representation separability on samples that have undergone superficial transformations to drive them out-of-distribution (OOD), such as typos or reformulations. By applying semantically-preserving perturbations, we study how separability degrades as statements become more OOD, across four LLM families, five evaluation datasets, and three knowledge probing methods. Our results reveal that internal representations of statement truthfulness collapse as the samples' presentations become less similar to those seen during pre-training. While LLMs can often distinguish between true and false statements when they closely resemble the pre-training data, this ability is highly dependent on the statement's exact surface form. These findings offer a possible explanation for brittle benchmark performance: LLMs may learn shallow, non-robust knowledge representations that allow for only limited generalizability. Our work presents a fundamental challenge for the utility of truthfulness probes, and more broadly, calls for further research on improving the robustness of learned knowledge representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。