专为中式菜品打造的高保真图像生成与编辑模型
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
- 构建首个针对中式菜品的图文生成框架,含最大规模数据集
- 通过分阶段训练与智能补全,显著提升细节还原度
- 支持精准编辑,适合餐饮数字化与电商视觉设计
菜肴图像在数字时代至关重要,随着食品产业与电商的数字化发展,对具有文化特色的菜肴图像需求持续增长。现有文本到图像生成模型虽能产出高质量图像,但在捕捉特定领域(尤其是中式菜肴)的多样特征与真实细节方面表现不足。为此,我们提出 Omni-Dish,首个专为中式菜肴设计的文本到图像生成模型。我们建立了一套完整的菜肴数据整理流程,构建了迄今最大的菜肴数据集。同时,引入重构描述策略并采用粗到细的训练方案,帮助模型更好地学习细微烹饪特征。推理时,利用预构建的高质量标题库与大语言模型增强用户文本输入,实现更逼真、更忠实的图像生成。此外,为扩展模型在菜肴编辑任务的能力,我们提出 Concept-Enhanced P2P,基于该方法构建了菜肴编辑数据集并训练专用编辑模型。大量实验验证了所提方法的优越性。
原文摘要 · Abstract (English)
Dish images play a crucial role in the digital era, with the demand for culturally distinctive dish images continuously increasing due to the digitization of the food industry and e-commerce. In general cases, existing text-to-image generation models excel in producing high-quality images; however, they struggle to capture diverse characteristics and faithful details of specific domains, particularly Chinese dishes. To address this limitation, we propose Omni-Dish, the first text-to-image generation model specifically tailored for Chinese dishes. We develop a comprehensive dish curation pipeline, building the largest dish dataset to date. Additionally, we introduce a recaption strategy and employ a coarse-to-fine training scheme to help the model better learn fine-grained culinary nuances. During inference, we enhance the user's textual input using a pre-constructed high-quality caption library and a large language model, enabling more photorealistic and faithful image generation. Furthermore, to extend our model's capability for dish editing tasks, we propose Concept-Enhanced P2P. Based on this approach, we build a dish editing dataset and train a specialized editing model. Extensive experiments demonstrate the superiority of our methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。