让AI按菜谱分步生成连贯图像,灵活适配不同长度步骤。
CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
- 分步区域控制:文本步骤与图像区域实时对齐。
- 生成多步图像序列,跨步骤食材保持一致。
- 适合教学视频、菜谱可视化等场景使用。
烹饪是具有顺序性和视觉依赖性的活动,每一步如切菜、混合或煎炸都包含程序逻辑和视觉语义。尽管最近的扩散模型在文本到图像生成方面表现出色,但在处理菜谱这类结构化多步场景时仍存在挑战。现有菜谱图像生成方法无法适应菜谱长度的自然变化,固定生成固定数量的图像。为此,我们提出 CookAnything,一个基于扩散模型的灵活且一致的框架,可从任意长度的文本烹饪指令中生成连贯、语义分明的图像序列。该框架引入三个关键组件:(1) 分步区域控制(SRC),在单个去噪过程中将文本步骤与对应图像区域对齐;(2) 灵活位置编码(Flexible RoPE),增强时间连贯性与空间多样性;(3) 跨步一致性控制(CSCC),保持各步骤间细粒度的食材一致性。在菜谱插图基准测试中,CookAnything 在有监督与无监督设置下均优于现有方法。该框架支持复杂多步指令的可扩展高质量视觉合成,具有广泛应用于教学媒体与程序化内容创作的潜力。
原文摘要 · Abstract (English)
Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusion models have shown strong capabilities in text-to-image generation, they struggle to handle structured multi-step scenarios like recipe illustration. Additionally, current recipe illustration methods are unable to adjust to the natural variability in recipe length, generating a fixed number of images regardless of the actual instructions structure. To address these limitations, we present CookAnything, a flexible and consistent diffusion-based framework that generates coherent, semantically distinct image sequences from textual cooking instructions of arbitrary length. The framework introduces three key components: (1) Step-wise Regional Control (SRC), which aligns textual steps with corresponding image regions within a single denoising process; (2) Flexible RoPE, a step-aware positional encoding mechanism that enhances both temporal coherence and spatial diversity; and (3) Cross-Step Consistency Control (CSCC), which maintains fine-grained ingredient consistency across steps. Experimental results on recipe illustration benchmarks show that CookAnything performs better than existing methods in training-based and training-free settings. The proposed framework supports scalable, high-quality visual synthesis of complex multi-step instructions and holds significant potential for broad applications in instructional media, and procedural content creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。