测试大模型真懂菜系文化,发现识别准但不会用文化知识
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

- 构建跨18地区10语言的菜系评测集,结合图像与烹饪流程
- 模型在标准题上超94%准确率,菜系归属题却降至56%以下
- 视觉无法激活已有知识,需显式连接过程与文化背景
多模态语言模型在食物识别基准上已接近完美表现,但其成功是源于真正的文化理解,还是仅靠视觉匹配尚不明确。为此,我们提出CulturalMenuBench,一个包含4,870个条目、覆盖10种语言和18个地区的基准,涵盖10项任务:将最终菜品图与步骤图、食材、流程文本及区域标签配对,从基础识别到基于过程的文化归因。评估12个模型发现显著的知识-应用鸿沟:尽管模型在标准多项选择任务上超过94%准确率,但在将菜肴归类至中国地域菜系时,准确率最高仅达56%,且格式相同。诊断分析显示:错误模式接近随机猜测,准确率与视觉差异相关,而非文化结构;仅凭菜名分类比依赖图像更准(+7–18分)。说明知识存在但无法通过视觉输入激活。消融实验表明,移除顺序烹饪图像会专门降低过程相关任务表现,其他任务保持稳定。总体而言,CulturalMenuBench揭示:高识别准确率可能掩盖文化知识应用能力缺失,呼吁训练中显式关联感知、流程与文化语境。代码与数据公开可用。
原文摘要 · Abstract (English)
Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。