用动态图分析优化大模型嵌入深度,提升心理测量结构发现效果
Optimizing the Landscape of LLM Embeddings with Dynamic Exploratory Graph Analysis for Generative Psychometrics: A Monte Carlo Study
- 将嵌入视为可探索的景观,用动态图分析逐层搜索最优维度范围
- 浅层嵌入(<300维)更准,深层(900-1298维)组织性更强但精度下降
- 综合指标比单一指标更优,嵌入深度随题项数量系统变化
大语言模型嵌入被用于心理量表项目池的维度结构预估,但现有方法将其视为静态截面表示,隐含假设所有嵌入坐标贡献均等,忽略了最优结构信息可能集中在特定空间区域。本研究将嵌入重构为可搜索的景观,采用动态探索图分析(DynEGA)系统遍历嵌入坐标,将维度索引作为伪时间序,类比密集纵向轨迹。基于OpenAI text-embedding-3-small模型,对代表五大自大型自恋维度的题项进行大规模蒙特卡洛模拟,系统改变题项池大小(每维度3–40题)和嵌入深度(3–1,298维)。结果表明:总熵拟合指数(TEFI)在深度900–1,200维时最小,此时熵组织性最强但结构准确性下降;归一化互信息(NMI)在浅层达到峰值,此时维度恢复最准确但熵拟合不佳。单一指标优化导致结构不一致,而加权复合准则识别出同时兼顾准确性和组织性的嵌入深度区间。最优嵌入深度随题项池规模系统变化。研究证明嵌入空间非均匀,需有原则地优化而非默认使用全向量。
原文摘要 · Abstract (English)
Large language model (LLM) embeddings are increasingly used to estimate dimensional structure in psychological item pools prior to data collection, yet current applications treat embeddings as static, cross-sectional representations. This approach implicitly assumes uniform contribution across all embedding coordinates and overlooks the possibility that optimal structural information may be concentrated in specific regions of the embedding space. This study reframes embeddings as searchable landscapes and adapts Dynamic Exploratory Graph Analysis (DynEGA) to systematically traverse embedding coordinates, treating the dimension index as a pseudo-temporal ordering analogous to intensive longitudinal trajectories. A large-scale Monte Carlo simulation embedded items representing five dimensions of grandiose narcissism using OpenAI's text-embedding-3-small model, generating network estimations across systematically varied item pool sizes (3-40 items per dimension) and embedding depths (3-1,298 dimensions). Results reveal that Total Entropy Fit Index (TEFI) and Normalized Mutual Information (NMI) leads to competing optimization trajectories across the embedding landscape. TEFI achieves minima at deep embedding ranges (900--1,200 dimensions) where entropy-based organization is maximal but structural accuracy degrades, whereas NMI peaks at shallow depths where dimensional recovery is strongest but entropy-based fit remains suboptimal. Single-metric optimization produces structurally incoherent solutions, whereas a weighted composite criterion identifies embedding dimensions depth regions that jointly balance accuracy and organization. Optimal embedding depth scales systematically with item pool size. These findings establish embedding landscapes as non-uniform semantic spaces requiring principled optimization rather than default full-vector usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。