用行为实验反推大模型语义结构,发现强制选择比自由联想更接近内部表征
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
- 通过强迫选择和自由联想实验,构建行为相似性矩阵
- 强制选择行为与隐藏层语义几何相关性显著高于自由联想(1750万次试验)
- 仅靠行为数据即可预测未见词的内部语义关系,适合认知建模研究
我们研究了大语言模型的隐藏状态几何结构在多大程度上可从其在心理语言学实验中的行为中恢复。在八个指令微调的Transformer模型上,使用两个实验范式——基于相似性的强迫选择和自由联想——在共享的5000词词汇集上进行了实验,收集超过1750万次试验,构建基于行为的相似性矩阵。通过表示相似性分析,将行为几何与各层隐藏状态相似性进行比较,并以FastText、BERT及跨模型共识为基准。结果表明,强迫选择行为与隐藏状态几何的相关性显著高于自由联想。在留出词的回归任务中,行为相似性(尤其是强迫选择)能超越词汇基线和跨模型共识,预测未见的隐藏状态相似性,说明仅凭行为测量就包含可恢复的内部语义几何信息。最后讨论了行为任务揭示隐藏认知状态的能力。
原文摘要 · Abstract (English)
We investigate the extent to which an LLM's hidden-state geometry can be recovered from its behavior in psycholinguistic experiments. Across eight instruction-tuned transformer models, we run two experimental paradigms -- similarity-based forced choice and free association -- over a shared 5,000-word vocabulary, collecting 17.5M+ trials to build behavior-based similarity matrices. Using representational similarity analysis, we compare behavioral geometries to layerwise hidden-state similarity and benchmark against FastText, BERT, and cross-model consensus. We find that forced-choice behavior aligns substantially more with hidden-state geometry than free association. In a held-out-words regression, behavioral similarity (especially forced choice) predicts unseen hidden-state similarities beyond lexical baselines and cross-model consensus, indicating that behavior-only measurements retain recoverable information about internal semantic geometry. Finally, we discuss implications for the ability of behavioral tasks to uncover hidden cognitive states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。