arXiv:2605.12074cs.CV2026-05

构建多任务咖啡制作视频基准,诊断模型在流程理解中的短板

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

论文配图:BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding
图 1 · 摘自论文原文
  • 基于185段第一视角视频构建细粒度场景图,关联物体、动作与步骤
  • 涵盖6类零样本语言任务,揭示不同任务间性能差异显著
  • 适合研究视频理解、具身智能和多模态推理的学者使用

场景理解是通用物理智能的核心,视频则能捕捉场景的状态与动态变化。然而,理解物理过程仍具挑战,因模型需融合物体定位、手物交互、关系解析、时间推理及步骤级过程推断。现有基准通常分项评估这些能力,限制了对模型失败原因的诊断。我们提出BARISTA,一个包含185段真实世界咖啡制作视频的密集标注第一视角数据集与基准,覆盖全自动、意式压粉和胶囊式三种流程。BARISTA提供每帧经验证的场景图,链接持久物体身份与掩码、轨迹、边界框、属性、类型化关系、手物交互、活动及流程步骤。基于这些图,我们构建了涵盖短语定位、手物交互识别、指代、活动识别、关系抽取和时序视觉问答的零样本语言任务。实验显示各任务族间表现差异显著,且无模型家族始终领先,表明BARISTA是极具挑战性的流程视频理解诊断基准。代码与数据集可在https://huggingface.co/datasets/ramblr/BARISTA获取。

原文摘要 · Abstract (English)

Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physical processes remains difficult, as models must combine object localization, hand-object interactions, relational parsing, temporal reasoning, and step-level procedural inference. Existing benchmarks usually evaluate these capabilities separately, limiting diagnosis of why models fail on procedural tasks. We introduce BARISTA, a densely annotated egocentric dataset and benchmark of 185 real-world coffee-preparation videos covering fully automatic, portafilter-based, and capsule-based workflows. BARISTA provides verified per-frame scene graphs linking persistent object identities to masks, tracks, boxes, attributes, typed relations, hand-object interactions, activities, and process steps. From these graphs, we derive zero-shot language-based tasks spanning phrase grounding, hand-object interaction recognition, referring, activity recognition, relation extraction, and temporal visual question answering. Experiments reveal strong variation across task families and no consistently dominant model family, positioning BARISTA as a challenging diagnostic benchmark for procedural video understanding. Code and dataset available at https://huggingface.co/datasets/ramblr/BARISTA.

视频理解多任务第一视角流程推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。