用视频分析方法研究印度不同地区咖喱饭的烹饪差异。
How Does India Cook Biryani?
- 构建12种地区风格的120个高清烹饪视频数据集,用视觉语言模型分割流程
- 通过音频、文本与视频对齐,自动识别并解释各地区做法差异
- 提出多层级问答评测基准,适合研究文化传承与多模态推理的学者
Biryani是印度最受推崇的菜肴之一,其制作方式、配料和呈现形式具有显著的区域多样性。随着在线烹饪视频的普及,利用计算工具系统性研究此类厨艺差异成为可能。然而,现有视频理解方法难以捕捉细粒度、多模态且根植于文化的烹饪视频差异。本文首次构建了一个大规模、精心整理的咖喱饭制作视频数据集,包含来自12个不同地区风格的120个高质量YouTube视频。我们提出一种多阶段框架,利用最新的视觉语言模型(VLMs)将视频分割为细粒度的烹饪步骤,并与音频转录和标准食谱文本对齐。基于这些对齐表示,我们引入一个视频比较管道,可自动识别并解释不同地区变体间的做法差异。我们构建了一个涵盖多个推理层级的问答(QA)基准,用于评估VLMs在程序性理解方面的能力。我们的方法采用多种VLM协同工作,结合人工验证确保高精度,并在零样本与微调设置下对比多个前沿模型。所提出的数据集、比较方法和问答基准为评估VLMs在结构化多模态推理任务中的表现提供了新平台,也为通过烹饪视频进行文化遗产的计算分析开辟了新方向。所有数据、代码及项目网站已公开:https://farzanashaju.github.io/how-does-india-cook-biryani/
原文摘要 · Abstract (English)
Biryani, one of India's most celebrated dishes, exhibits remarkable regional diversity in its preparation, ingredients, and presentation. With the growing availability of online cooking videos, there is unprecedented potential to study such culinary variations using computational tools systematically. However, existing video understanding methods fail to capture the fine-grained, multimodal, and culturally grounded differences in procedural cooking videos. This work presents the first large-scale, curated dataset of biryani preparation videos, comprising 120 high-quality YouTube recordings across 12 distinct regional styles. We propose a multi-stage framework leveraging recent advances in vision-language models (VLMs) to segment videos into fine-grained procedural units and align them with audio transcripts and canonical recipe text. Building on these aligned representations, we introduce a video comparison pipeline that automatically identifies and explains procedural differences between regional variants. We construct a comprehensive question-answer (QA) benchmark spanning multiple reasoning levels to evaluate procedural understanding in VLMs. Our approach employs multiple VLMs in complementary roles, incorporates human-in-the-loop verification for high-precision tasks, and benchmarks several state-of-the-art models under zero-shot and fine-tuned settings. The resulting dataset, comparison methodology, and QA benchmark provide a new testbed for evaluating VLMs on structured, multimodal reasoning tasks and open new directions for computational analysis of cultural heritage through cooking videos. We release all data, code, and the project website at https://farzanashaju.github.io/how-does-india-cook-biryani/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。