arXiv:2604.25423cs.CLcs.AI2026-04ACL

用指示词测试大模型是否理解身体认知和文化差异

Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives

论文配图:Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives
图 1 · 摘自论文原文
  • 以英汉语的'这/那'为探针,检验模型对空间指代的理解
  • 五款主流大模型均无法区分远近指代,且无文化差异表现
  • 适合研究语言认知、模型偏差与跨文化人工智能的读者

大型语言模型(LLMs)是否真正从文本中习得具身认知与文化惯例?我们引入指示词——如英语中的'this/that'和中文的'这/那'——作为探测具身知识的新工具。基于320名母语者提供的6,400条回应,我们建立人类基准:英语使用者能可靠区分近距远距指代,但视角转换能力弱;中文使用者则能灵活切换视角,容忍远距模糊。相比之下,五款顶尖大模型未能内在理解近距远距差异,且未表现出文化差异,始终呈现英语中心化推理。本研究贡献:(i) 提出以指示词为基础的新评估任务,用于检验具身认知与文化惯例;(ii) 提供人类跨文化解读不对称性的实证证据;(iii) 为自我中心-社会中心之争提供新视角,揭示两种取向共存但因语言而异;(iv) 呼吁未来模型设计应关注个体差异。

原文摘要 · Abstract (English)

Do large language models (LLMs) truly acquire embodied cognition and cultural conventions from text? We introduce demonstratives, fundamental spatial expressions like "this/that" in English and "zhè/nà" in Chinese, as a novel probe for grounded knowledge. Using 6,400 responses from 320 native speakers, we establish a human baseline: English speakers reliably distinguish proximal-distal referents but struggle with perspective-taking, while Chinese speakers switch perspectives fluently but tolerate distal ambiguity. In contrast, five state-of-the-art LLMs fail to inherently understand the proximal-distal contrast and show no cultural differences, defaulting to English-centric reasoning. Our study contributes (i) a new task, based on demonstratives, as a new lens for evaluating embodied cognition and cultural conventions; (ii) empirical evidence of cross-cultural asymmetries in human interpretation; (iii) a new perspective on the egocentric-sociocentric debate, showing both orientations coexist but vary across languages; and (iv) a call to address individual variation in future model design.

具身认知文化差异语言模型指示词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。