arXiv:2501.09042cs.CVcs.GR2025-01被引 7

用扩散模型生成连贯的烹饪步骤图像,让虚拟厨房更真实。

CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion

  • 引入记忆网络融合文本、图像和多模态提示,生成连贯烹饪序列。
  • 在YouCookII数据集上达到高保真度与步骤一致性,FID值显著降低。
  • 支持食材与做法修改,适合智能烹饪助手与数字内容创作。

近期文本到图像生成模型在创造多样化且逼真的图像方面取得显著进展,这一成果已延伸至食品图像生成领域,通过烹饪风格、食材和食谱等条件输入实现多样输出。然而,基于食谱中的烹饪步骤生成一系列连贯的程序化图像仍是未被探索的挑战。这有望提升烹饪体验的视觉引导效果,并推动智能烹饪模拟系统的构建。为此,我们提出一项新任务——烹饪程序化图像生成。该任务极具挑战性,需生成与烹饪步骤一致且具有时序一致性的逼真图像。为此,我们提出CookingDiffusion,利用Stable Diffusion并结合三种创新的记忆网络(Memory Nets)来建模程序化提示,包括文本提示(对应烹饪步骤)、图像提示(对应烹饪图像)以及多模态提示(混合步骤与图像),以确保生成图像的连续性。为验证方法有效性,我们对YouCookII数据集进行预处理,建立新基准。实验结果表明,该模型在生成高质量烹饪序列图像方面表现优异,时序一致性指标显著提升,同时在FID和提出的平均程序一致性(Average Procedure Consistency)评估中均优于现有方法。此外,CookingDiffusion具备修改食谱中食材与烹饪方式的能力。代码、模型与数据集将公开共享。

原文摘要 · Abstract (English)

Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called \textbf{cooking procedural image generation}. This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while preserving sequential consistency. To collectively tackle these challenges, we present \textbf{CookingDiffusion}, a novel approach that leverages Stable Diffusion and three innovative Memory Nets to model procedural prompts. These prompts encompass text prompts (representing cooking steps), image prompts (corresponding to cooking images), and multi-modal prompts (mixing cooking steps and images), ensuring the consistent generation of cooking procedural images. To validate the effectiveness of our approach, we preprocess the YouCookII dataset, establishing a new benchmark. Our experimental results demonstrate that our model excels at generating high-quality cooking procedural images with remarkable consistency across sequential cooking steps, as measured by both the FID and the proposed Average Procedure Consistency metrics. Furthermore, CookingDiffusion demonstrates the ability to manipulate ingredients and cooking methods in a recipe. We will make our code, models, and dataset publicly accessible.

图像生成扩散模型烹饪模拟多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。