arXiv:2507.21529cs.CV2025-07中稿 · ACM MM 2025被引 4

用双向思维链引导,让每步菜谱生成更连贯准确的图像。

Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance

  • 通过动态选图块参考历史生成结果,确保食材外观匹配描述。
  • 在多个数据集上显著提升图像连贯性与语义一致性,最高提升23.6%。
  • 适合需要可视化烹饪过程的研究者或智能食谱系统开发者。

烹饪过程可视化是图像生成与食品分析交叉领域的一项有前景的任务,旨在为食谱中的每个步骤生成对应的图像。然而,现有方法多聚焦于基于食谱生成完成菜品的图像,难以实现烹饪过程的可视化,面临两大挑战:一是食材在不同步骤中外观变化多样,难以生成与文本描述一致的正确外观,导致语义不一致;二是当前步骤可能依赖前序操作,必须保持图像序列的上下文连贯性。为此,本文提出一种名为 Chain-of-Cooking 的烹饪过程可视化模型。为生成准确的食材外观,设计了动态图块选择模块(Dynamic Patch Selection Module),从先前生成的图像块中检索与当前文本内容最相关的参考信息。为增强序列连贯性,提出语义演化模块(Semantic Evolution Module)和双向思维链(Bidirectional Chain-of-Thought, CoT)引导机制。语义演化模块建立潜在提示与当前烹饪步骤之间的语义关联,并将其融合至潜在特征;双向思维链则更新融合后的特征,以指导当前步骤保持与前序步骤的一致性。此外,构建了一个名为 CookViz 的新数据集,包含烹饪过程的中间图文对。定量与定性实验表明,该方法在生成连贯且语义一致的烹饪过程图像方面优于现有方法。

原文摘要 · Abstract (English)

Cooking process visualization is a promising task in the intersection of image generation and food analysis, which aims to generate an image for each cooking step of a recipe. However, most existing works focus on generating images of finished foods based on the given recipes, and face two challenges to visualize the cooking process. First, the appearance of ingredients changes variously across cooking steps, it is difficult to generate the correct appearances of foods that match the textual description, leading to semantic inconsistency. Second, the current step might depend on the operations of previous step, it is crucial to maintain the contextual coherence of images in sequential order. In this work, we present a cooking process visualization model, called Chain-of-Cooking. Specifically, to generate correct appearances of ingredients, we present a Dynamic Patch Selection Module to retrieve previously generated image patches as references, which are most related to current textual contents. Furthermore, to enhance the coherence and keep the rational order of generated images, we propose a Semantic Evolution Module and a Bidirectional Chain-of-Thought (CoT) Guidance. To better utilize the semantics of previous texts, the Semantic Evolution Module establishes the semantical association between latent prompts and current cooking step, and merges it with the latent features. Then the CoT Guidance updates the merged features to guide the current cooking step remain coherent with the previous step. Moreover, we construct a dataset named CookViz, consisting of intermediate image-text pairs for the cooking process. Quantitative and qualitative experiments show that our method outperforms existing methods in generating coherent and semantic consistent cooking process.

图像生成烹饪可视化思维链视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。