线性关系越强,大模型越容易编造虚假答案。
Relational Linearity is a Predictor of Hallucinations
- 用线性嵌入机制解释幻觉:模型易为未知主体生成合理但虚构的客体。
- 在15种关系上测试,线性关系与幻觉率相关性高达0.58至0.84。
- 适合关注模型可靠性与幻觉控制的研究者参考。
幻觉是语言模型的核心缺陷。本文研究模型对合成实体的回应,例如“格伦·古尔德演奏过什么乐器?”——这些实体对模型而言完全未知。我们发现,如Gemma-7B-IT等模型常产生幻觉,难以识别虚构事实不在其知识范围内。基于线性关系嵌入的思路,提出假设:(i) 由于抽象表示方式,模型能轻易为非真实主体生成看似合理的对象,导致幻觉;(ii) 非线性关系无此生成机制,更难幻觉。为此构建了包含15种关系的合成未知实体基准数据集SyntHal。在四个指令微调模型上验证,关系线性程度与幻觉倾向高度相关,相关系数 $r ext{ in } [.58, .84]$。
原文摘要 · Abstract (English)
Hallucination is a central failure mode of language models (LMs). We focus on hallucinations in response to questions like: "Which instrument did Glenn Gould play?", but we ask these questions for synthetic entities designed to be unknown to the model. We find that LMs like Gemma-7B-IT frequently hallucinate, i.e., they have difficulty recognizing that the hallucinated fact is not part of their knowledge. Based on the idea of linear relational embeddings, we put forward the following hypothesis. (i) Due to the abstract scheme that is used to represent them, LMs can easily produce plausible objects for non-existing subjects of linear relations, which can lead to hallucinations. (ii) For a nonlinear relation, this mechanism for producing an object is not available and so a hallucination is easier to avoid. To test this hypothesis, we create SyntHal, a synthetic unknown-entity benchmark for 15 relations. We find that across four instruction-tuned models, relational linearity is a strong predictor of models hallucinating an object for an unknown subject vs refusing to give an answer, with correlations $r \in [.58, .84]$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。