arXiv:2509.05566cs.CLcs.CY2025-09

人类新词约定能泛化到未讨论对象,说明语言有深层概念协同。

Ad hoc conventions generalize to new referents

  • 通过重复沟通建立临时命名规则,测试其对新图像的泛化能力。
  • 配对者对未讨论图像的描述一致性显著提升,且与视觉相似度非线性相关。
  • 结果支持语言是概念协同而非随意标签,适合研究语言习得与智能体设计。

人们如何谈论从未提及的事物?一种观点认为新命名系统是任意关联特定目标的,如同人名无法扩展;另一种观点认为共享描述方式涉及更广的概念对齐,重塑个体语义空间,应能泛化到新对象。我们在一项双人通信实验中(N=302)使用最新发布的KiloGram数据集(含1,000多张抽象拼图图像),让参与者先就一组图像达成指称惯例,再评估其对未讨论图像的描述一致性。结果显示:配对者相对于初始标签,对新图像的描述一致性显著提升。该泛化效应随视觉相似度呈非线性下降(符合谢泼德定律),且在不同可命名性水平下均稳健。这些发现表明,临时约定并非随意标签,而是真实的概念协调,对参考理论和语言智能体设计具有启示。

原文摘要 · Abstract (English)

How do people talk about things they've never talked about before? One view suggests that a new shared naming system establishes an arbitrary link to a specific target, like proper names that cannot extend beyond their bearers. An alternative view proposes that forming a shared way of describing objects involves broader conceptual alignment, reshaping each individual's semantic space in ways that should generalize to new referents. We test these competing accounts in a dyadic communication study (N=302) leveraging the recently-released KiloGram dataset containing over 1,000 abstract tangram images. After pairs of participants coordinated on referential conventions for one set of images through repeated communication, we measured the extent to which their descriptions aligned for undiscussed images. We found strong evidence for generalization: partners showed increased alignment relative to their pre-test labels. Generalization also decayed nonlinearly with visual similarity (consistent with Shepard's law) and was robust across levels of the images' nameability. These findings suggest that ad hoc conventions are not arbitrary labels but reflect genuine conceptual coordination, with implications for theories of reference and the design of more adaptive language agents.

语言习得概念对齐泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。