研究智能体如何从视觉经验中习得并稳定新词汇,发现感知距离决定学习难易。
Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents

- 用冻结的DINOv2视觉特征和假词测试语言学习,模拟真实语义习得过程。
- 感知距离越近,词汇学习越准;语义相关性对学习无显著影响。
- 命名与检索能力不同,记忆保真度是独立于命名准确性的关键维度。
当前人工智能系统多以任务表现和行为模仿评估,但未检验其能否从具身经验中获得、稳定并使用新词汇意义。本文提出『词汇共识』实验框架,基于冻结的DINOv2视觉嵌入、卡罗尔风格的假词,以及可解释的词汇学习器与线性基线,测试智能体是否能为视觉概念习得人工标签,实现双向泛化,并在受控条件下保持稳定。主要结果呈现稳健的感知一致性梯度:原生类别最易学习,一致泛化仍可习得,中等离散概念性能下降,远距离离散概念接近随机水平。预注册的CIFAR-100分离实验确认,该梯度由感知距离主导而非语义关联:感知距离可预测学习准确率(部分R² = 0.245,p < 1e-7),而语义距离无显著解释力(部分R² = 0.002,p = 0.660)。双向评估显示命名与检索是独立能力:基于实例的机制在标签到图像检索中优于中心原型,揭示了独立于命名准确性的记忆保真维度。伪造控制、同质候选池评估及表征重构的零结果表明,冻结的感知几何结构既支持词汇具身化,也限制了无表征适配下的学习范围。
原文摘要 · Abstract (English)
Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such evaluations leave open whether an artificial agent can acquire, stabilize, and use new lexical meanings from grounded experience. This paper introduces Lexical Consensus, an experimental framework for studying grounded word learning over a structured perceptual substrate. Using frozen DINOv2 visual embeddings, Carroll-style nonce words, and interpretable lexical learners plus linear baselines, we test whether agents can acquire artificial labels for visual concepts, generalize them bidirectionally, and stabilize them across controlled settings. The main result is a robust perceptual-coherence gradient: native categories are easiest to learn, coherent overextensions remain learnable, mid-range disjunctive concepts degrade, and far-disjunctive concepts approach chance. A pre-registered CIFAR-100 dissociation experiment confirms that this gradient is governed by perceptual distance rather than semantic relatedness: perceptual distance predicts acquisition accuracy (partial R^2 = 0.245, p < 1e-7), while semantic distance adds no significant explanatory power (partial R^2 = 0.002, p = 0.660). Bidirectional evaluation shows that naming and retrieval are distinct: exemplar-based mechanisms outperform centroid prototypes in label-to-image retrieval, exposing a memory-fidelity dimension separate from naming accuracy. Falsification controls, homogeneous candidate-pool evaluations, and null results on representational restructuring indicate that frozen perceptual geometry both enables lexical grounding and limits what can be acquired without representational adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。