arXiv:2607.17559cs.ROcs.AI2026-07

让机器人通过气味、视觉和语言理解环境,实现嗅觉定位。

COLIP-2: Olfaction-Vision-Language Embeddings

论文配图:COLIP-2: Olfaction-Vision-Language Embeddings
图 1 · 摘自论文原文
  • 将分子结构、传感器数据、语言描述与图像统一到同一表征空间。
  • 首次实现机器人根据气味概率定位场景中的物体。
  • 优化后可在边缘设备实时运行,适合机器人应用。

对比嗅觉-语言-图像预训练2(COLIP-2)模型构建了一个多模态嵌入空间,将嗅觉置于与视觉和语言同等重要的位置。分子结构、气体传感器读数、气味描述语言及图像均被训练至单一共享表征空间,使机器人能基于检测到的气味概率性地定位场景中的物体。目前尚无大规模配对图像-气味数据集,因此需推动新数据集的建立。本文旨在展示开源嗅觉数据所能实现的极限,以论证发展新方法与数据集对实现先进嗅觉感知能力的必要性。我们报告了内部测试结果,并对模型进行必要优化,使其可在边缘设备实时运行,适用于实际机器人应用。尽管为机器人设计,但其架构受到学术界与工业界多领域专家影响,期望在任何需要嗅觉智能的多模态领域发挥作用。

原文摘要 · Abstract (English)

The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that places olfaction as a first-class citizen among vision and language. Molecular structure, gas-sensor readings, odor-descriptor language, and images are all trained into a single shared representation space, so that a robot can localize a detected aroma to objects in a scene probabilistically. No ImageNet-scale datasets of paired image-scent examples exists which warrants the need for their collection. Our intent with the release of COLIP-2 is to demonstrate the limit of what can be built for robotics with open-sourced olfactory data in order to ground the argument for why new methodologies and datasets are necessary in order to enable advanced olfactory-oriented perception capabilities. We enumerate results from internal testing of the COLIP-2 architecture and make necessary optimizations to run the model at the edge for real-time robotics applications. While developed with robotics in mind, the design of COLIP-2 has been influenced by experts across many disciplines of science in academia and industry, and we hope that the model can be useful in any multimodal domain requiring olfactory intelligence.

嗅觉感知多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。