比较人类与大模型在语义空间中的搜索方式,发现模型无法完全模拟人类的探索平衡。
Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing
- 用轨迹分析法量化语义搜索的三维度:熵、步长、离心率。
- 人类搜索更随机、步幅更大、覆盖更广,模型均不及。
- 调温度只能局部匹配,无配置能复现人类完整行为模式。
语义记忆检索可视为在概念空间中的导航。本研究通过言语流畅性数据,对比了人类与三个大语言模型(GPT-4o、Gemini-2.5-Pro、Claude-Sonnet-4.5)的语义搜索动态。对82名人类参与者及大模型在八种温度设置下的生成项,采用基于轨迹的自然语言处理指标,量化了三个互补维度:熵(步长可预测性)、到下一跳距离(连续语义步长)、到中心点距离(全局分散度)。结果表明,人类在熵值、语义步长和全局分散度上均高于所有大模型,显示其搜索更具变异性与探索性。温度调节仅实现部分对齐:个别指标在特定设置下与人类接近,但无任何配置能同时匹配全部三个维度的人类特征。这说明当前模型架构尚无法再现人类语义搜索中局部利用与全局探索之间的独特平衡。
原文摘要 · Abstract (English)
Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency data. By applying trajectory-based NLP metrics to the items generated by 82 human participants and LLM output across eight temperature settings, we quantified three complementary dimensions: entropy (step size predictability), distance to next (successive semantic steps), and distance to centroid (global dispersion). Humans exhibited higher entropy, larger semantic steps and broader dispersion than all LLMs, indicating more variable and exploratory search. Temperature tuning produced only partial alignments, as individual metrics matched between humans and LLMs at specific settings, but no configuration reproduced the complete human profile (in all dimensions). These findings suggest that human semantic search implements a distinctive balance between local exploitation and global exploration that current model architectures fail to reproduce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。