用损失描述恢复大模型认知不确定性的本质差异
From Scalars to Tensors: Declared Losses Recover Epistemic Distinctions That Neutrosophic Scalars Cannot Express

- 将标量三元组扩展为带损失声明的张量结构,捕捉模型真实认知状态
- 84%的模型在无约束评估中出现超真值现象(T+I+F>1.0)
- 损失词汇差异显著区分悖论、无知与偶然等不同不确定性类型
Leyva-Vázquez 和 Smarandache(2025)证明了中和性三值评价(真、不确定、假独立且不强制和为1.0)可揭示35%复杂认知情形中的“超真”(T+I+F > 1.0)现象。我们在此基础上拓展:第一,在五家厂商(Anthropic、Meta、DeepSeek、Alibaba、Mistral)共五个模型族中复现并扩展实验,发现84%的无约束评估存在超真,证实该现象跨厂商普遍;第二,识别出标量三值表达的局限:当模型采取“吸收”立场(T=0, I=1, F=0)时,悖论、无知、偶然等根本不同的认知情境输出相同标量,导致语义差异被掩盖。我们进一步证明,引入声明损失(结构化描述模型无法评估及其原因)可有效恢复这些区分。产生相同标量的模型在损失关键词上的杰卡德相似度低于0.10,其领域特异、程度分级的损失声明能准确反映不确定性的本质差异。这表明标量三值仅为必要条件,而张量结构(标量+损失)才是更真实的认知状态表征。
原文摘要 · Abstract (English)
Leyva-Vázquez and Smarandache (2025) demonstrated that neutrosophic T/I/F evaluation, where Truth, Indeterminacy, and Falsity are independent dimensions not constrained to sum to 1.0, which reveals "hyper-truth"' (T+I+F > 1.0) in 35% of complex epistemic cases evaluated by LLMs. We extend their work in two directions. First, we replicate and extend their experiment across five model families from five vendors (Anthropic, Meta, DeepSeek, Alibaba, Mistral), finding hyper-truth in 84% of unconstrained evaluations, which confirms the phenomenon is cross-vendor under our prompt protocol. Second, and more significantly, we identify a limitation of scalar T/I/F that their framework cannot address: models adopting an `"Absorption" position (T=0, I=1, F=0) produce identical scalar outputs for fundamentally different epistemic situations (paradox, ignorance, contingency), collapsing the very distinctions neutrosophic logic was designed to preserve. We demonstrate that extending the evaluation to include declared losses (structured descriptions of what the model cannot evaluate and why) substantially recovers these distinctions. Models producing identical scalars for paradox and ignorance produce nearly disjoint loss vocabularies (Jaccard similarity < 0.10 on loss description keywords), with domain-specific, severity-rated loss declarations that differentiate the nature of their uncertainty. This suggests that scalar T/I/F is a necessary but insufficient representation of epistemic state, and that tensor-structured output (scalars + losses) provides a more faithful model of LLM epistemic capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。