通过识别食材状态提升视障者做饭进度追踪精度
Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking
- 基于食材工具状态变化构建烹饪步骤识别框架
- 在173段视频和12次真实场景中提升步骤预测准确率
- 适合无障碍烹饪系统设计与视障用户辅助研究
烹饪对日常独立与健康至关重要,但视障者因缺乏进度追踪与情境反馈支持而面临挑战。物体状态(如食材或工具的形态变化)为上下文感知的烹饪辅助提供了潜在基础,却未被充分探索。本文提出OSCAR(Object Status Context Awareness for Recipes)技术流程,通过解析食谱、提取物体状态、对齐视觉与烹饪步骤、时间因果建模,实现非视觉烹饪中的实时步骤追踪。在173段教学视频及12次由盲人用户在家中录制的真实烹饪会话数据集上评估,结果表明物体状态能持续提升多模态模型的步骤预测准确率,并揭示了隐式任务、摄像头位置与光照条件等影响实际表现的关键因素。研究贡献包括可复用的上下文感知进度追踪流程、标注的真实世界非视觉烹饪数据集,以及未来辅助烹饪系统的设计洞见。
原文摘要 · Abstract (English)
Cooking plays a vital role in everyday independence and well-being, yet remains challenging for people with vision impairments due to limited support for tracking progress and receiving contextual feedback. Object status - the condition or transformation of ingredients and tools - offers a promising but underexplored foundation for context-aware cooking support. In this paper, we present OSCAR (Object Status Context Awareness for Recipes), a technical pipeline that explores the use of object status recognition to enable recipe progress tracking in non-visual cooking. OSCAR integrates recipe parsing, object status extraction, visual alignment with cooking steps, and time-causal modeling to support real-time step tracking. We evaluate OSCAR on 173 instructional videos and a real-world dataset of 12 non-visual cooking sessions recorded by BLV individuals in their homes. Our results show that object status consistently improves step prediction accuracy across vision-language models, and reveal key factors that impact performance in real-world conditions, such as implicit tasks, camera placement, and lighting. We contribute the pipeline of context-aware recipe progress tracking, an annotated real-world non-visual cooking dataset, and design insights to guide future context-aware assistive cooking systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。