arXiv:2509.00033cs.CVcs.AI2025-09

用多模态模型分析厨房动作,自动生成烹饪步骤指南。

Deep Learning-Driven Multimodal Detection and Movement Analysis of Objects in Culinary

  • 融合YOLOv8、LSTM与Whisper的多模态系统识别厨具与动作。
  • 基于手部点序列与语音识别数据,驱动小模型生成完整菜谱。
  • 专为复杂厨房环境设计,适合智能厨房与自动化助手应用。

本研究探索现有模型并进行微调,结合YOLOv8分割模型、基于手部关键点运动序列训练的LSTM模型,以及语音识别模型Whisper-base,提取足够数据供小型语言模型TinyLLaMa预测菜谱并生成分步操作文本。所有数据均由作者采集,构建了一个针对复杂挑战性环境优化的任务专用系统,验证了计算机视觉在日常活动如烹饪中的可扩展性与广泛应用潜力。该工作拓展了计算机视觉在日常生活任务中的边界。

原文摘要 · Abstract (English)

This is a research exploring existing models and fine tuning them to combine a YOLOv8 segmentation model, a LSTM model trained on hand point motion sequence and a ASR (whisper-base) to extract enough data for a LLM (TinyLLaMa) to predict the recipe and generate text creating a step by step guide for the cooking procedure. All the data were gathered by the author for a robust task specific system to perform best in complex and challenging environments proving the extension and endless application of computer vision in daily activities such as kitchen work. This work extends the field for many more crucial task of our day to day life.

厨房视觉多模态动作分析智能烹饪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。