arXiv:2604.00002cs.CLcs.AI2026-04被引 4

测试大模型对气味的推理能力,发现它们更依赖词汇关联而非分子结构分析。

Benchmark for Assessing Olfactory Perception of Large Language Models

论文配图:Benchmark for Assessing Olfactory Perception of Large Language Models
图 1 · 摘自论文原文
  • 构建8类1010个气味问题的基准,涵盖分类、描述词识别等任务
  • 使用化合物名提示比同分异构体SMILES提示平均高7个百分点,最高提升18.9%
  • 多语言集成模型在气味预测上达到0.86的AUROC,适合跨语言研究者

本文提出嗅觉感知(OP)基准,用于评估大语言模型(LLM)对气味的推理能力。该基准包含1010个问题,覆盖八类任务:气味分类、主要描述词识别、强度与愉悦度判断、多描述词预测、混合物相似性、嗅觉受体激活及真实气味源识别。每个问题以化合物名称和同分异构体SMILES两种格式呈现,以评估分子表示方式的影响。在21种主流模型配置上评估发现,化合物名提示始终优于SMILES提示,性能提升2.4至18.9个百分点(均值约7个百分点),表明当前LLM主要通过词汇关联而非结构化分子推理获取嗅觉知识。表现最佳模型总体准确率达64.4%,揭示了嗅觉推理能力的初步显现与显著差距。进一步在21种语言上评估子集,发现多语言预测集成可提升性能,最佳语言集成模型的AUROC达0.86。大模型应能处理嗅觉信息,而不仅限于视觉或听觉。

原文摘要 · Abstract (English)

Here we introduce the Olfactory Perception (OP) benchmark, designed to assess the capability of large language models (LLMs) to reason about smell. The benchmark contains 1,010 questions across eight task categories spanning odor classification, odor primary descriptor identification, intensity and pleasantness judgments, multi-descriptor prediction, mixture similarity, olfactory receptor activation, and smell identification from real-world odor sources. Each question is presented in two prompt formats, compound names and isomeric SMILES, to evaluate the effect of molecular representations. Evaluating 21 model configurations across major model families, we find that compound-name prompts consistently outperform isomeric SMILES, with gains ranging from +2.4 to +18.9 percentage points (mean approx +7 points), suggesting current LLMs access olfactory knowledge primarily through lexical associations rather than structural molecular reasoning. The best-performing model reaches 64.4\% overall accuracy, which highlights both emerging capabilities and substantial remaining gaps in olfactory reasoning. We further evaluate a subset of the OP across 21 languages and find that aggregating predictions across languages improves olfactory prediction, with AUROC = 0.86 for the best performing language ensemble model. LLMs should be able to handle olfactory and not just visual or aural information.

大模型评测嗅觉认知多语言分子表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。