语言模型能提前判断对实体的熟悉程度,且可调控拒绝回答行为。
Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering
- 通过分析模型最后一层激活值,检测对实体的熟悉度。
- 波兰语适配模型的熟悉度评分与实体流行度高度相关(ρ=0.28-0.57)。
- 可调节数值控制模型拒绝回答,适合安全生成场景使用。
语言模型能否在生成答案前评估对某个实体的熟悉程度?我们研究了来自Bielik、PLLuM、Gemma-4和Qwen3系列的12个指令微调模型,在包含1440个波兰实体的新数据集上,覆盖四个领域和十个维基页面浏览量分位数,以及虚构对照组。所有模型家族中,熟悉度探测分数均能区分真实与虚构实体;在波兰语适配的Bielik和PLLuM系列中,该分数还与实体流行度显著相关(模型平均斯皮尔曼ρ为0.28–0.57),而Gemma-4和Qwen3系列最多仅为0.11,这一模式更与波兰语适配相关而非参数量。在双模型对比实验中,当波兰语提问句替换为英文但实体名不变时,探测器在同语言下的AUROC保持96%–101%,显示对提示语言的鲁棒性。在Gemma-4-12B中,唯一原生支持拒绝的模型,仅在一个层添加一维熟悉度方向即可单调调节拒绝率(知名实体从0.24升至1.00,未知实体从0.73降至0.00)。最后,校准后的熟悉度探测器在生成前放弃机制中表现良好,但生成后检测器在预测行为错误方面平均更优。这些结果支持存在分级的生成前实体熟悉度读出,并表明表征熟悉度与将其转化为拒绝策略的政策是分离的。
原文摘要 · Abstract (English)
Can a language model estimate its familiarity with an entity before generating an answer? We study activations at the final prompt token in twelve instruction-tuned models from the Bielik, PLLuM, Gemma-4, and Qwen3 families, using a new dataset of 1,440 Polish entities spanning four domains and ten Wikipedia-pageview deciles, plus fabricated controls. Familiarity-probe scores separate real from fabricated entities in every family; in the Polish-adapted Bielik and PLLuM families they additionally track entity popularity (model-mean Spearman $ρ$ 0.28-0.57, versus at most 0.11 in Gemma-4 and Qwen3), a pattern more strongly associated with Polish adaptation than with parameter count in this model sample. In a paired experiment on two families, probes retain 96-101% of within-language AUROC when the Polish question stem is replaced with an English one around unchanged entity names, showing robustness to prompt language in this setting. In Gemma-4-12B, the only model that natively refuses, adding a one-dimensional familiarity direction at a single layer moves refusal rates monotonically in both directions (0.24 to 1.00 on well-known entities; 0.73 to 0.00 on unknown ones). Finally, a calibrated familiarity probe is competitive among pre-generation abstention gates, although post-generation detectors better predict behavioral error on average. These results support a graded pre-generation entity-familiarity readout, and a separation between representational familiarity and the policy that converts it into abstention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。