用掩码精修生成烹饪动作视觉辅助图,环境一致更真实。
VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
- 通过掩码定位动作对象,精准修改局部画面。
- 在三个视频数据集上优于当前最优方法。
- 适合需要实时视觉指导的智能厨具研发者。
烹饪不仅需要遵循步骤,还需理解、执行和监控每一步,缺乏视觉引导时难度较大。尽管食谱图片和视频提供线索,但其焦点、工具和布局常不一致。为更好支持烹饪过程,我们提出VisualChef,一种生成情境化视觉辅助图像的方法。给定初始帧和指定动作,VisualChef生成展示动作执行过程及目标物体结果状态的图像,同时保持初始环境一致性。以往工作依赖大语言模型提取知识并生成详细文本描述来引导图像生成,需精细图文对齐且依赖额外标注。相比之下,VisualChef通过掩码实现视觉定位,简化对齐过程。核心思路是识别动作相关对象并分类,实现针对性修改以反映预期动作与结果,同时维持环境连贯性。此外,我们设计自动化流程提取高质量初始、动作和最终状态帧。我们在三个第一人称视频数据集上进行定量与定性评估,结果表明其性能优于现有先进方法。
原文摘要 · Abstract (English)
Cooking requires not only following instructions but also understanding, executing, and monitoring each step - a process that can be challenging without visual guidance. Although recipe images and videos offer helpful cues, they often lack consistency in focus, tools, and setup. To better support the cooking process, we introduce VisualChef, a method for generating contextual visual aids tailored to cooking scenarios. Given an initial frame and a specified action, VisualChef generates images depicting both the action's execution and the resulting appearance of the object, while preserving the initial frame's environment. Previous work aims to integrate knowledge extracted from large language models by generating detailed textual descriptions to guide image generation, which requires fine-grained visual-textual alignment and involves additional annotations. In contrast, VisualChef simplifies alignment through mask-based visual grounding. Our key insight is identifying action-relevant objects and classifying them to enable targeted modifications that reflect the intended action and outcome while maintaining a consistent environment. In addition, we propose an automated pipeline to extract high-quality initial, action, and final state frames. We evaluate VisualChef quantitatively and qualitatively on three egocentric video datasets and show its improvements over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。