arXiv:2606.27378cs.CLcs.LG2026-06

提出四条公理评估大模型隐式思维表征,揭示基准测试掩盖的表征缺陷。

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

论文配图:Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
图 1 · 摘自论文原文
  • 基于因果、最小性、可分性和稳定性四条公理,独立量化表征质量。
  • 23个推理任务中无模型同时满足四条公理,表征无法区分同任务内不同问题。
  • 表征信息基本来自输入嵌入,结构缺陷普遍存在,与模型规模和训练方式无关。

我们提出一个大模型隐式思维表征的公理化评估框架,包含独立于下游基准得分的度量,能揭示基准准确率掩盖的表征缺陷。现有评估将表征质量与模型能力混淆,导致失败归因不清。我们形式化四个功能公理(因果性、最小性、可分性、稳定性),并定义可直接在表征上计算的定量度量。我们在23个推理任务(如空间推理、事实问答)中审计了开源大模型。结果发现:没有模型同时满足全部四条公理;表征能可靠区分任务类型,但无法区分同一任务内的不同问题;表征所含信息几乎完全源自输入嵌入。该缺陷在密集型、推理蒸馏型及强化学习训练的模型族中均一致存在,表明其为结构性问题,而非模型大小或训练方式所致。

原文摘要 · Abstract (English)

We introduce an axiomatic evaluation framework for latent thought representations in LLMs, comprising metrics that are independent of downstream benchmark scores and reveal representational failures that benchmark accuracy masks. Existing evaluations conflate representation quality with model capacity. Therefore, failures cannot be attributed to the representation rather than to the model that processes it. We formalize four functional axioms (Causality, Minimality, Separability, and Stability) and define a quantitative measure for each, computed directly on the representation independently of downstream accuracy. We audit open-weight LLMs across 23 reasoning tasks (e.g., Spatial Reasoning, Factual QA). We find that no candidate satisfies all four axioms simultaneously, that the representations distinguish task type reliably but cannot distinguish between two questions within the same task, and that the representations encode little information beyond what is already present in the input embedding. The failure is consistent across dense, reasoning-distilled, and RL-trained model families, indicating that the gap is structural rather than a property of model size or training procedure.

大模型表征分析推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。