一个能同时处理图文菜谱生成的通用食品模型。
ChefFusion: Multimodal Foundation Model Integrating Recipe and Food Image Generation
- 融合语言与图像模型,统一处理多种食品任务。
- 在图文生成任务上表现优于已有模型。
- 适合研究食品多模态生成的开发者使用。
食品计算领域已有大量研究,但多数仅聚焦单一任务,如从菜名和食材生成食谱(t2t)、从食物图像生成食谱(i2t),或从食谱生成食物图像(t2i)。现有方法未实现多模态的协同处理。为此,我们提出一种新型食品计算基础模型ChefFusion,具备真正多模态能力,可同时支持t2t、t2i、i2t、it2t及t2ti等任务。该模型结合大语言模型(LLM)与预训练图像编码器/解码器,实现食品理解、识别、食谱生成与图像生成等多样化功能。相比先前模型,本模型展现出更广的任务覆盖范围与更强性能,尤其在食谱与图像生成任务中表现突出。代码已开源至GitHub。
原文摘要 · Abstract (English)
Significant work has been conducted in the domain of food computing, yet these studies typically focus on single tasks such as t2t (instruction generation from food titles and ingredients), i2t (recipe generation from food images), or t2i (food image generation from recipes). None of these approaches integrate all modalities simultaneously. To address this gap, we introduce a novel food computing foundation model that achieves true multimodality, encompassing tasks such as t2t, t2i, i2t, it2t, and t2ti. By leveraging large language models (LLMs) and pre-trained image encoder and decoder models, our model can perform a diverse array of food computing-related tasks, including food understanding, food recognition, recipe generation, and food image generation. Compared to previous models, our foundation model demonstrates a significantly broader range of capabilities and exhibits superior performance, particularly in food image generation and recipe generation tasks. We open-sourced ChefFusion at GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。