构建首个跨感官食物体验数据集,从图片预测味觉、嗅觉等多维度感受。
FoodSense: A Multisensory Food Dataset and Benchmark for Predicting Taste, Smell, Texture, and Sound from Images

- 基于66,842组人标注数据,建立跨感官感知的图文对应关系
- 模型可同时输出感官评分与视觉化解释,支持多维感知预测
- 适合作为多模态模型在食品感知任务中的基准评测工具
人类常通过食物图像推断味道、气味、质感甚至声音,这一现象在认知科学中已有研究。然而,以往的食物视觉语言研究主要集中在识别任务,如餐食分类、成分检测和营养估算,对图像驱动的多感官体验预测仍缺乏探索。本文提出FoodSense,一个由人工标注的跨感官推理数据集,包含66,842个参与者-图像配对,覆盖2,987张独特食物图像。每对数据包含1-5分的数值评分及自由文本描述,涵盖味道、气味、质感和声音四个感官维度。为支持模型生成可解释的预测,将简短的人类注释扩展为基于图像的推理链条,使用大语言模型生成条件化的视觉理由。基于这些标注,训练出FoodSense-VL模型,可直接从食物图像生成多感官评分与可解释的推理解释。本研究连接认知科学中的跨感官感知发现与现代指令微调技术,揭示现有主流评估指标在视觉感官推断任务中的不足。
原文摘要 · Abstract (English)
Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has focused primarily on recognition tasks such as meal identification, ingredient detection, and nutrition estimation. Image-based prediction of multisensory experience remains largely unexplored. We introduce FoodSense, a human-annotated dataset for cross-sensory inference containing 66,842 participant-image pairs across 2,987 unique food images. Each pair includes numeric ratings (1-5) and free-text descriptors for four sensory dimensions: taste, smell, texture, and sound. To enable models to both predict and explain sensory expectations, we expand short human annotations into image-grounded reasoning traces. A large language model generates visual justifications conditioned on the image, ratings, and descriptors. Using these annotations, we train FoodSense-VL, a vision language benchmark model to produce both multisensory ratings and grounded explanations directly from food images. This work connects cognitive science findings on cross-sensory perception with modern instruction tuning for multimodal models and shows that many popular evaluation metrics are insufficient for visually sensory inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。