LLM嵌入能还原心理健康领域的专家分类结构,但效果依赖层级和需控制干扰因素。
Do LLM Embedding Spaces Recover Expert Structure?

- 用专家症状矩阵对比嵌入空间,评估其几何结构是否匹配专业分类
- 微调后在细粒度类别上对齐专家结构更显著,大模型规模提升零样本与微调效果
- 即使控制情感、语言风格等干扰,仍保留显著对齐,适合心理语言学研究
预训练文本嵌入被广泛用作表征地图,但高类别可分性不等于几何结构恢复了专家定义的模式。我们以心理健康相关语言为研究对象,利用28个Reddit社区数据,在0.6B和4B规模的Qwen3模型上比较预训练与监督微调嵌入空间。构建类别原型,通过表示相似性分析(RSA)将其与专家症状矩阵对比,并辅以原型典型性评估及多基线干扰控制。结果表明:预训练嵌入在心理健康子集内已具备可观测对齐;微调显著增强细粒度类别上的对齐;大模型规模同时提升零样本对齐与微调增益。在控制情感(VAD)、LIWC、词汇风格及话题分布结构后,残余对齐依然显著。说明LLM嵌入可恢复专家相关的类别几何结构,但该恢复具有层级依赖性,应通过显式干扰控制验证,而非仅依赖分类性能推断。
原文摘要 · Abstract (English)
Pretrained text embeddings are increasingly used as representational maps, yet high category separability does not imply that their geometry recovers expert-defined structure. We study this problem in mental-health-related language, where symptom relations provide an external reference and online communities introduce strong domain, affective, stylistic, and discourse confounds. Using 28 Reddit communities, we compare pretrained and supervised fine-tuned Qwen3 embedding spaces at two scales (0.6B and 4B). We construct category prototypes, evaluate their representational dissimilarity matrices against an expert symptom matrix with representational similarity analysis, and complement this global test with prototype-based typicality and multi-baseline confound controls. Pretrained embeddings show measurable alignment with expert structure within the mental-health subset; fine-tuning strengthens this alignment most at the finest category level; and larger scale improves both zero-shot alignment and supervision-induced gains. Residual alignment remains substantial after controlling for VAD, LIWC, lexical style, and topic-distribution structure. These results suggest that LLM embeddings can recover expert-relevant category geometry, but this recovery is level-dependent and should be tested against explicit confounds rather than inferred from classification alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。