arXiv:2503.05962cs.HCcs.CV2025-03被引 12

用物体状态追踪帮视障者做饭,实时提供上下文反馈。

OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking

  • 结合大模型与视觉语言模型,动态追踪食材状态变化
  • 相比基线方法提升超20%的进度识别准确率
  • 适合无障碍烹饪辅助系统研发与残障科技应用

视障人士在烹饪过程中遵循食谱是一项重要但困难的任务。我们开发了OSCAR(Recipe对象状态与上下文感知系统),通过追踪物体状态来实现食谱进度跟踪和任务完成的上下文感知反馈。OSCAR利用大型语言模型(LLMs)和视觉语言模型(VLMs),对食谱步骤进行操作、提取物体状态信息、将视觉帧与物体状态对齐,并生成烹饪进度日志。我们使用173个YouTube烹饪视频和12个真实世界非视觉烹饪视频评估了OSCAR的食谱跟随能力,结果表明,相比基线方法,使用物体状态可使性能提升超过20%,并分析了影响预测表现的关键因素。此外,我们还贡献了一个包含步标注的真实非视觉烹饪视频数据集,作为评估基准。

原文摘要 · Abstract (English)

Following recipes while cooking is an important but difficult task for visually impaired individuals. We developed OSCAR (Object Status Context Awareness for Recipes), a novel approach that provides recipe progress tracking and context-aware feedback on the completion of cooking tasks through tracking object statuses. OSCAR leverages both Large-Language Models (LLMs) and Vision-Language Models (VLMs) to manipulate recipe steps, extract object status information, align visual frames with object status, and provide cooking progress tracking log. We evaluated OSCAR's recipe following functionality using 173 YouTube cooking videos and 12 real-world non-visual cooking videos to demonstrate OSCAR's capability to track cooking steps and provide contextual guidance. Our results highlight the effectiveness of using object status to improve performance compared to baseline by over 20% across different VLMs, and we present factors that impact prediction performance. Furthermore, we contribute a dataset of real-world non-visual cooking videos with step annotations as an evaluation benchmark.

无障碍技术视觉语言模型烹饪辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。