arXiv:2512.09095cs.CV2025-12中稿 · WACV 2026

解决多名词食物生成中成分混淆问题,提升真实感。

Food Image Generation on Multi-Noun Categories

  • 引入食物领域知识,早期优化多词关系理解
  • 在UEC-256上显著减少错误成分出现,提升布局准确性
  • 适合图像生成、餐饮AI等场景应用

为包含多个名词的食物类别生成真实图像极具挑战性。例如,提示词‘蛋面’可能生成鸡蛋和面条作为独立实体的图像。多名词食物类别在真实数据集广泛存在,占UEC-256等基准数据集的很大比例。此类复合名称常导致生成模型误解语义,产生不恰当的食材或物体。这源于文本编码器中多名词类别知识不足及对多名词关系的误读,进而引发空间布局错误。为此,我们提出FoCULR(Food Category Understanding and Layout Refinement),融合食物领域知识,并在生成过程早期引入核心概念。实验表明,这些技术的整合显著提升了食品领域的图像生成性能。

原文摘要 · Abstract (English)

Generating realistic food images for categories with multiple nouns is surprisingly challenging. For instance, the prompt "egg noodle" may result in images that incorrectly contain both eggs and noodles as separate entities. Multi-noun food categories are common in real-world datasets and account for a large portion of entries in benchmarks such as UEC-256. These compound names often cause generative models to misinterpret the semantics, producing unintended ingredients or objects. This is due to insufficient multi-noun category related knowledge in the text encoder and misinterpretation of multi-noun relationships, leading to incorrect spatial layouts. To overcome these challenges, we propose FoCULR (Food Category Understanding and Layout Refinement) which incorporates food domain knowledge and introduces core concepts early in the generation process. Experimental results demonstrate that the integration of these techniques improves image generation performance in the food domain.

图像生成多名词理解食物生成布局优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。