模型对实体的熟悉度可通过激活分布提前判断,且与回答可靠性相关。
Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale

- 用激活分散度分析模型在生成答案前对实体的熟悉程度。
- 1.5B到11B模型均能以0.95以上准确率区分真实与虚构实体。
- 该信号早于答案生成,适合用于检测模型是否‘知道’自己不知道。
大语言模型在面对从未见过的实体时最容易幻觉。我们探究模型在生成首个回答词之前,其激活是否能揭示实体熟悉度,以及这一信号是否预测回答的事实可靠性。在四个波兰语BieliK模型(1.5B-11B参数)上,针对运动员、城市、作家、音乐家四类实体(每类42个知名、42个冷门真实、42个虚构),使用单句提问(每模型504个提示)。通过后SwiGLU MLP层的两个无监督、单前向传播的分散度指标(逆参与比率和谱熵),在所有领域与规模下均实现0.95-1.00的AUROC,监督线性探针达0.99-1.00。两者均显著优于随机置换基线(0.70-0.74,p≤1e-3),在保留层选择下仍稳定(0.93-0.99),真实名称区分效果良好(已知vs.冷门真实:0.96-1.00)。信号跨实体类型迁移(对角线外平均AUROC 0.92-0.99);对照实验表明,主要下降源于模板影响而非实体类型。该信号在头间广泛分布。表示信号在1.5B时已达上限,而行为层面的事实可靠性随规模急剧上升:1.5B、4.5B、7B、11B模型分别正确回答42个已知运动员中的0、2、10、19个(严格评分)。对于已知实体,区分正确与幻觉回答较难(探针0.93;分散度不优于首词熵基线)。五样本语义熵基线需5倍推理成本,仅达0.71-0.83。尽管存在内部认知,模型几乎从不拒绝或回避:审计2520个回答中仅2次拒绝、1次含糊。实体熟悉度与事实可靠性是不同尺度下的独立现象。
原文摘要 · Abstract (English)
Large language models hallucinate most about entities they have never seen. We ask whether a model's activations betray entity familiarity before a single answer token is generated, and whether that signal predicts the factual reliability of the answers. On four Polish Bielik models (1.5B-11B parameters), we probe four entity domains (athletes, cities, writers, musicians), each with 42 well-known, 42 obscure-but-real, and 42 fabricated entities addressed by a one-sentence question (504 prompts per model). Two unsupervised, single-forward-pass dispersion measures over post-SwiGLU MLP activations, inverse participation ratio and spectral entropy, separate known from fabricated entities at AUROC 0.95-1.00 across all domains and scales; a supervised linear probe reaches 0.99-1.00. Both clear selection-aware permutation floors of about 0.70-0.74 (empirical p<=1e-3), survive held-out layer selection (0.93-0.99), and persist on real names (known vs. obscure-but-real: 0.96-1.00). The signal transfers across entity types (mean off-diagonal AUROC 0.92-0.99); a matched-template counterfactual shows the only large drops are template-caused, not entity-type effects, and the signal is diffuse across heads. This representational signal is already at ceiling at 1.5B, whereas behavioral factual reliability scales sharply: 0, 2, 10, and 19 of 42 known athletes are answered fully correctly by the 1.5B, 4.5B, 7B, and 11B models under a strict judge. Within known entities, separating correct from hallucinated answers is much harder (probe 0.93; dispersion no better than a first-token-entropy baseline). A five-sample semantic-entropy baseline reaches only 0.71-0.83 at 5x the inference cost. Despite this internal awareness, the models almost never abstain: an audit of 2,520 answers finds 2 refusals and 1 hedge. Entity familiarity and factual reliability are distinct phenomena on different scaling curves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。