无需训练,通过提示嫁接实现多食物图像精准生成
Training-Free Text-to-Image Compositional Food Generation via Prompt Grafting
- 用文本显式空间提示+采样时隐式布局引导,分两阶段生成
- 在两个食物数据集上显著提升目标物体出现率,分离效果可控
- 适合饮食评估、食谱可视化等需要多食材精确生成的场景
现实中的餐食图像通常包含多个食物成分,这对基于图像的膳食评估和食谱可视化等应用中多食材数据增强提出了需求。然而,当前主流文本到图像扩散模型因物体纠缠问题(如米饭与汤混合)难以准确生成多食材图像,主要因为许多食物缺乏清晰边界。为此,我们提出无需训练的提示嫁接(Prompt Grafting, PG)框架,结合文本中的显式空间线索与采样过程中的隐式布局引导。该方法采用两阶段流程:先通过布局提示建立独立区域,待布局稳定后再嫁接目标提示。此框架实现了食物纠缠的可控性:用户可通过编辑布局排列,指定哪些食物应分离或故意混合。在两个食物数据集上的实验表明,该方法显著提升了目标物体的生成率,并提供了可控分离的定性证据。
原文摘要 · Abstract (English)
Real-world meal images often contain multiple food items, making reliable compositional food image generation important for applications such as image-based dietary assessment, where multi-food data augmentation is needed, and recipe visualization. However, modern text-to-image diffusion models struggle to generate accurate multi-food images due to object entanglement, where adjacent foods (e.g., rice and soup) fuse together because many foods do not have clear boundaries. To address this challenge, we introduce Prompt Grafting (PG), a training-free framework that combines explicit spatial cues in text with implicit layout guidance during sampling. PG runs a two-stage process where a layout prompt first establishes distinct regions and the target prompt is grafted once layout formation stabilizes. The framework enables food entanglement control: users can specify which food items should remain separated or be intentionally mixed by editing the arrangement of layouts. Across two food datasets, our method significantly improves the presence of target objects and provides qualitative evidence of controllable separation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。