让大模型在未知环境自创共享词汇,用视觉与语义对齐实现跨智能体沟通。
Lexical discovery in unknown environments orchestrated by Large Language Models

- 基于大模型的智能体通过参照游戏自组织出共享的外星词汇体系。
- 20个智能体面对10个新视觉对象,达成共识且收敛性模型拟合度超0.95。
- 发现的词锚定于人类语言嵌入空间,适合深空或深海自主探索任务。
在未知环境(如行星或深海探测)中部署的自主智能体群体必须为无名实体创建共享词汇。本文提出神经符号词汇发现(NSLD)框架:基于大语言模型的智能体在分布外视觉参照物上进行指称游戏,自主构建共享的外星词汇体系。每个智能体结合冻结的CLIP视觉编码器、私有FAISS向量索引和仅文本的LLM。关键在于,所发现的外星词汇通过嵌入空间中的语义相近性锚定于自然语言,从而将新的感知基底词汇扩展至人类词汇库。在最多20个智能体与10个视觉参照物的仿真中实现共识。通过三种分析模型刻画收敛动态,拟合优度R² > 0.95,为自主探测任务的预部署规划迈出第一步。
原文摘要 · Abstract (English)
Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies to refer to entities that have no name in any human language. We propose the Neuro-Symbolic Lexical Discovery (NSLD) framework, in which a population of LLM-based agents plays a referential game over out-of-distribution visual referents, autonomously self-organising a shared alien lexicon. Each agent combines a frozen CLIP vision encoder with a private FAISS vector index and a text-only LLM. Crucially, discovered alien words are anchored to natural language via semantic proximity in the embedding space, enlarging the human vocabulary with new perceptually grounded words. Consensus is reached in simulations with populations of up to twenty agents and ten visual referents. Convergence dynamics are characterised through three analytical models achieving R^2 > 0.95, representing a first step towards pre-deployment planning in autonomous exploration missions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。