用自然语言描述实现无人集群去中心化人员识别,更透明易懂。
Language-Based Swarm Perception: Decentralized Person Re-Identification via Natural Language Descriptions
- 每台机器人用视觉语言模型生成人物外观的文本描述
- 通过文本比对与聚类,无需中心协调即可识别同一人
- 支持语言查询和可解释性,适合需要透明决策的场景
本文提出一种在机器人集群中去中心化进行人员再识别的方法,以自然语言为主要表征方式。不同于依赖高维视觉嵌入(即图像提取的特征向量)的传统方法,该方法让每个机器人本地使用视觉语言模型(VLM)检测并生成人物外观的可读文本描述,而非特征向量。这些文本描述在集群中分布式比对与聚类,无需中央协调即可将同一人的观测归为一组。随后由语言模型对每个聚类提炼出代表性描述,形成群体感知的简洁、可解释摘要。该方法支持自然语言查询,提升透明度,促进可解释的集群行为。初步实验表明,在身份一致性与可解释性方面表现优于或接近基于嵌入的方法,尽管当前存在文本相似度精度与计算开销的限制。后续工作将优化相似度度量、探索语义导航,并拓展语言感知至环境要素。本研究强调去中心化感知与通信,主动导航仍是未来开放方向。
原文摘要 · Abstract (English)
We introduce a method for decentralized person re-identification in robot swarms that leverages natural language as the primary representational modality. Unlike traditional approaches that rely on opaque visual embeddings -- high-dimensional feature vectors extracted from images -- the proposed method uses human-readable language to represent observations. Each robot locally detects and describes individuals using a vision-language model (VLM), producing textual descriptions of appearance instead of feature vectors. These descriptions are compared and clustered across the swarm without centralized coordination, allowing robots to collaboratively group observations of the same individual. Each cluster is distilled into a representative description by a language model, providing an interpretable, concise summary of the swarm's collective perception. This approach enables natural-language querying, enhances transparency, and supports explainable swarm behavior. Preliminary experiments demonstrate competitive performance in identity consistency and interpretability compared to embedding-based methods, despite current limitations in text similarity and computational load. Ongoing work explores refined similarity metrics, semantic navigation, and the extension of language-based perception to environmental elements. This work prioritizes decentralized perception and communication, while active navigation remains an open direction for future study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。