arXiv:2608.12852cs.CLcs.AI2026-08

AI把不可能和虚假混为一谈,但内部表征其实能区分两者。

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

论文配图:Falsehood and Impossibility Are Different Directions in an AI's Representation of Language
图 1 · 摘自论文原文
  • 通过激活分析发现模型内部可区分'不可能'与'虚假'。
  • 第15层的不可能性探测器准确率达97%,对矛盾句识别差。
  • 不可能性与语义异常在表征上接近,但可区分,适合哲学与AI认知研究者。

语言可描述虚假情况与根本不可能的情况。人工智能模型是否在内部区分这两类失败尚不明确。本文对多模态开源模型Gemma 3 4B IT进行探索性激活研究,使用85个来自17个哲学主题的提示,每个主题包含真陈述、偶然错误、非可能陈述、语义异常和必然错误五类。模型在回答中将12个错误陈述标记为‘矛盾’,但其激活模式显示不同:线性真值探针能有效分离不可能与真实陈述(AUC 0.93),却无法区分不可能与虚假(AUC 0.20)。在保留主题族上的不可能性探针达到AUC 1.00,第15层准确率97%(Bonferroni校正后P=0.018)。真值与不可能性方向近乎正交,而不可能性方向部分重叠语义异常方向但可区分。稀疏自编码器特征在同层重现该几何结构,仅不可能性特征在异常句上活跃,极少响应偶然错误句。因此,模型激活空间中,必然错误并非偶然错误的极端形式,而是更接近实验定义的语义异常类别。这一相关性观察虽不意味着不可能陈述本体无意义,但为古老哲学区分提供了实证注脚。

原文摘要 · Abstract (English)

Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood. In its answers, the model conflates contingent falsehood with contradiction, labeling 12 of 15 false statements "contradiction." Its activations show a different pattern. A linear truth probe separates impossible from true statements (AUC 0.93) but not impossible from false statements (AUC 0.20). An impossibility probe evaluated on held-out topic families separates necessary from contingent falsehood at AUC 1.00, peaking at layer 15 with balanced accuracy 0.97 (Bonferroni-adjusted P=0.018). The truth and impossibility directions are close to orthogonal, whereas the impossibility direction partially overlaps a semantic anomaly direction while remaining distinguishable from it. Sparse autoencoder features at the same layer repeat this geometry. Features selective for impossibility also fire on anomalous sentences but rarely on contingent falsehoods. In this model's activation space, necessary falsehoods are not extreme cases of contingent falsehood but lie closer to the experimentally defined category of semantic anomaly. This representational proximity does not imply that impossible statements are intrinsically meaningless. These correlational observations from one small model offer an empirical footnote to an old philosophical distinction.

语言表征逻辑推理神经机制哲学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。