用视觉图像自动生成嗅觉数据,实现跨模态嗅觉感知新范式
See & Sniff: Learning Visuo-Olfactory Representations

- 通过语义一致的图像合成嗅觉样本,构建大规模跨模态数据集
- 自监督学习联合表征,在嗅觉分类上比纯嗅觉模型高7%
- 首次提出像素级嗅觉定位任务,适合多模态感知与机器人研究者
当前多模态模型主要融合视觉与语言、音频或触觉,但嗅觉因缺乏配对的视嗅数据而未被充分探索。本文提出SmellNet-V,基于同一语义类别内气味身份对视觉变换不变的特性,将无配对的纯嗅觉样本与真实网络图像进行语义对齐,无需耗时的同步采集即可构建跨模态基准。在此基础上,我们提出See & Sniff框架,通过密集局部对齐学习联合视嗅表征,并自然生成气味显著性图以实现气味源的空间定位。进一步引入像素级嗅觉定位任务及评估基准。实验表明,该方法在仅凭嗅觉输入的分类任务中较纯嗅觉基线提升7%,且可泛化至跨模态检索与气味定位任务,确立了视嗅学习作为多模态感知的新方向。
原文摘要 · Abstract (English)
While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows us to synthetically pair smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection. Building on this dataset, we propose See & Sniff, a self-supervised framework that learns joint visuo-olfactory representations via dense local alignment and naturally produces smell saliency maps for spatial grounding of odor sources. We further introduce pixel-level smell localization task and a benchmark for evaluation. Our method surpasses smell-only baselines by 7% in smell classification from smell alone and generalizes to cross-modal retrieval and smell localization, establishing visuo-olfactory learning as a new direction in multimodal perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。