arXiv:2602.05971cs.CLcs.LG2026-02中稿 · ICLR被引 1

将人类概念生成过程建模为嵌入空间中的轨迹,量化语义导航行为。

Characterizing Human Semantic Navigation in Concept Production as Trajectories in Embedding Space

  • 用累积嵌入构建个体语义轨迹,提取几何与动力学特征。
  • 在多语言任务中区分临床群体与概念类型,效果优于传统方法。
  • 框架可适配不同模型与语言,适合认知研究与人工认知评估。

语义表征可视为人类在其中检索与操作意义的结构化动态知识空间。为探究人类如何在此几何空间中导航,我们提出一种将概念生成建模为嵌入空间中导航的框架。基于不同Transformer文本嵌入模型,利用累积嵌入构建参与者特定的语义轨迹,并提取距离下一概念、距中心点距离、熵、速度、加速度等几何与动力学指标。这些度量捕捉了语义导航的标量与方向性特征,为语义表征搜索提供了计算基础。我们在四个跨语言数据集上验证该框架:神经退行性疾病、骂人词流畅性、意大利语属性列举任务及德语任务。结果表明,该方法能有效区分临床组别与概念类型,且相比传统人工语言预处理,所需人力极少。对比非累积方法发现,累积嵌入在长轨迹中表现更优,而短轨迹因上下文不足更适合非累积方案。值得注意的是,不同嵌入模型结果相似,表明尽管训练路径不同,其学习表征仍具共性。通过将语义导航建模为嵌入空间中的结构化轨迹,本研究连接认知建模与学习表征,建立了量化语义动态的可扩展流程,适用于临床研究、跨语言分析与人工认知评估。

原文摘要 · Abstract (English)

Semantic representations can be framed as a structured, dynamic knowledge space through which humans navigate to retrieve and manipulate meaning. To investigate how humans traverse this geometry, we introduce a framework that represents concept production as navigation through embedding space. Using different transformer text embedding models, we construct participant-specific semantic trajectories based on cumulative embeddings and extract geometric and dynamical metrics, including distance to next, distance to centroid, entropy, velocity, and acceleration. These measures capture both scalar and directional aspects of semantic navigation, providing a computationally grounded view of semantic representation search as movement in a geometric space. We evaluate the framework on four datasets across different languages, spanning different property generation tasks: Neurodegenerative, Swear verbal fluency, Property listing task in Italian, and in German. Across these contexts, our approach distinguishes between clinical groups and concept types, offering a mathematical framework that requires minimal human intervention compared to typical labor-intensive linguistic pre-processing methods. Comparison with a non-cumulative approach reveals that cumulative embeddings work best for longer trajectories, whereas shorter ones may provide too little context, favoring the non-cumulative alternative. Critically, different embedding models yielded similar results, highlighting similarities between different learned representations despite different training pipelines. By framing semantic navigation as a structured trajectory through embedding space, bridging cognitive modeling with learned representation, thereby establishing a pipeline for quantifying semantic representation dynamics with applications in clinical research, cross-linguistic analysis, and the assessment of artificial cognition.

语义导航嵌入空间认知建模跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。