arXiv:2603.12117cs.CLcs.AI2026-03

用葡萄酒品鉴测试大模型的感官判断能力,发现文本学习难抵真实味觉经验。

SommBench: Assessing Sommelier Expertise of Language Models

  • 构建多语言葡萄酒知识测评基准,含理论问答、风味补全和搭配推荐三任务
  • 顶尖模型在理论题上正确率达97%,但风味补全仅65%,搭配任务准确率不足40%
  • 专为检验语言模型能否通过文字模拟专家感官判断而设计,适合研究具身认知的学者

随着大语言模型快速发展,系统评估其多语言与跨文化能力变得日益重要。现有文化评测基准主要关注可由语言编码的基础文化知识。本文提出SommBench,一个用于评估语言模型品酒专家能力的多语言基准,该领域深度依赖嗅觉与味觉感知。尽管语言模型仅通过文本描述学习感官属性,SommBench旨在检验这种文本表征是否足以模拟专家级感官判断。SommBench包含三项核心任务:葡萄酒理论问答(WTQA)、葡萄酒特征补全(WFC)和食物-葡萄酒搭配(FWP),支持英语、斯洛伐克语、瑞典语、芬兰语、德语、丹麦语、意大利语及西班牙语。数据集由专业侍酒师与各语种母语者共同开发,共包含1,024道葡萄酒理论问答题、1,000个风味补全样例和1,000个食物搭配样例。我们报告了主流语言模型的表现,包括Gemini 2.5等闭源模型以及GPT-OSS和Qwen 3等开源模型。结果显示,最先进模型在理论问答任务中表现优异(最高97%正确率),但在风味补全任务中仅达65%峰值,食物搭配任务的马修相关系数(MCC)在0至0.39之间,表明此类任务更具挑战性。SommBench为评估语言模型品酒能力提供了新颖且具有挑战性的评测平台,数据集已公开于https://github.com/sommify/sommbench。

原文摘要 · Abstract (English)

With the rapid advances of large language models, it becomes increasingly important to systematically evaluate their multilingual and multicultural capabilities. Previous cultural evaluation benchmarks focus mainly on basic cultural knowledge that can be encoded in linguistic form. Here, we propose SommBench, a multilingual benchmark to assess sommelier expertise, a domain deeply grounded in the senses of smell and taste. While language models learn about sensory properties exclusively through textual descriptions, SommBench tests whether this textual grounding is sufficient to emulate expert-level sensory judgment. SommBench comprises three main tasks: Wine Theory Question Answering (WTQA), Wine Feature Completion (WFC), and Food-Wine Pairing (FWP). SommBench is available in multiple languages: English, Slovak, Swedish, Finnish, German, Danish, Italian, and Spanish. This helps separate a language model's wine expertise from its language skills. The benchmark datasets were developed in close collaboration with a professional sommelier and native speakers of the respective languages, resulting in 1,024 wine theory question-answering questions, 1,000 wine feature-completion examples, and 1,000 food-wine pairing examples. We provide results for the most popular language models, including closed-weights models such as Gemini 2.5, and open-weights models, such as GPT-OSS and Qwen 3. Our results show that the most capable models perform well on wine theory question answering (up to 97% correct with a closed-weights model), yet feature completion (peaking at 65%) and food-wine pairing show (MCC ranging between 0 and 0.39) turn out to be more challenging. These results position SommBench as an interesting and challenging benchmark for evaluating the sommelier expertise of language models. The benchmark is publicly available at https://github.com/sommify/sommbench.

大模型评测多语言感官认知侍酒师

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。