发现Transformer模型能按心理状态连续分布编码,形成类似光谱的结构。
Probing Spectrum-Like Organization of States of Mind in Transformer Representation Spaces
- 用636条语句构建心理状态分级数据集,标注连续分值与7级有序标签。
- 五种预训练模型均能准确识别连续分数和离散层级,性能显著优于随机基线。
- 几何分析显示状态分布呈低到高有序排列,适合研究心智状态的表征机制。
我们探究了心理状态是否在Transformer表示空间中形成类似光谱的连续结构。为此,构建了一个包含636个短自然语言句子的数据集,每条句子同时标注了从-5到5的连续评分以及7个有序层级(从僵化、匮乏表达到更连贯、反思性、整合性表达)。评估了五种冻结的Transformer表示:四种句子嵌入模型和一种解码器仅有的残差流表示。在所有表示中,简单探测器可靠地恢复了连续评分和离散层级标签,置换测试表明性能显著高于打乱标签的基线。进一步分析揭示了一致的几何模式:UMAP投影显示从低到高的有序组织,混淆矩阵误差集中于相邻层级之间,方向性消融识别出一个显著对齐评分的成分。这些结果表明,Transformer表示中存在统计显著的、与标注的心理状态结构一致的谱状组织。标注仅作为表示分析的操作框架,不具临床或诊断意义。
原文摘要 · Abstract (English)
We investigate whether graded states of mind form spectrum-like structure in transformer representation spaces. To do so, we construct a dataset of 636 short natural-language sentences annotated with both a continuous score from $-5$ to $5$ and one of seven ordered tiers, ranging from collapsed or scarcity-driven expressions to more coherent, reflective, and integrative ones. We evaluate five frozen transformer representations: four sentence-embedding models and one decoder-only residual-stream representation. Across all representations, simple probes reliably recover both the continuous score and the discrete tier labels, and permutation tests show that performance significantly exceeds shuffled-label baselines. Additional analyses reveal a consistent geometric pattern: UMAP projections show low-to-high organization, confusion matrices concentrate errors between neighboring tiers, and directional ablation identifies a prominent score-aligned component. These results suggest that transformer representations contain statistically significant, spectrum-like organization aligned with the annotated state-of-mind structure. The annotations are used only as an operational framework for representation analysis, not as a clinical or diagnostic measure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。