arXiv:2607.06402cs.CVcs.AI2026-07

用语言连接视觉与嗅觉,让模型理解场景气味来源。

What Images Cannot Say: Language-Guided Olfactory Representation Learning

论文配图:What Images Cannot Say: Language-Guided Olfactory Representation Learning
图 1 · 摘自论文原文
  • 用视觉语言模型生成场景描述,引导嗅觉表征学习。
  • 在纽约气味数据集上,嗅觉到图像检索性能超越基线。
  • 可分离物体与环境的气味成分,结果可解释。

图像告诉我们场景看起来什么样,却很少说明身临其境的感受。尽管已有数据集将视觉场景与电子鼻测量值配对,但将气味信号与图像对齐仍具挑战性,因为许多嗅觉线索源于不可见的环境上下文。我们提出SCENT,一种多模态框架,利用语言作为视觉与嗅觉之间的语义桥梁。该方法借助视觉-语言模型(VLMs)生成包含物体、环境上下文及视觉场景暗示的潜在气味线索的场景描述,为嗅觉表征学习提供语义指导。我们训练了一个气味编码器,将电子鼻信号映射到与视觉和文本表示对齐的共享嵌入空间,并引入语言引导的潜在分解,以分离物体特定气味与环境贡献。在纽约气味数据集上的实验表明,相比仅使用视觉的基线方法,SCENT显著提升了跨模态检索性能,在气味到图像和气味到文本检索任务中达到当前最优水平。此外,该框架生成的嗅觉表征具有可解释性,能实现复杂气味混合物的解耦。结果揭示了上下文语义信息在多模态学习中对嗅觉感知建模的重要性,为该领域未来发展奠定基础。

原文摘要 · Abstract (English)

Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose measurements, aligning smell signals with images remains challenging because many olfactory cues arise from contextual environmental factors that are not directly visible in pixels. We introduce SCENT, a multimodal framework that uses language guidance as a semantic bridge between vision and olfaction. Our approach leverages Vision-Language Models (VLMs) to generate scene descriptors capturing objects, environmental context, and plausible ambient smell cues suggested by the visual scene. These descriptors provide semantic guidance for learning olfactory representations. We train a smell encoder that maps electronic-nose signals into a shared embedding space aligned with both visual and textual representations, and introduce a languageguided latent decomposition that separates object-specific odors from contextual environmental contributions. Experiments on the New York Smells dataset demonstrate that SCENT significantly improves crossmodal retrieval compared to vision-only baselines, achieving state-of-theart performance on smell-to-image and smell-to-text retrieval tasks. In addition, our framework produces interpretable olfactory representations that enable the disentanglement of complex smell mixtures. Our results reveal the importance of contextual semantic information for grounding olfactory perception in multimodal learning and pave the way for future research in this area.

多模态嗅觉表征语言引导可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。