arXiv:2409.18459cs.CVcs.MM2024-09被引 3

用日式菜谱微调多模态大模型,生成食材更准

FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation

  • 在日式菜谱数据上微调开源多模态模型
  • 食材生成F1达0.531,优于GPT-4o的0.481
  • 适合做饮食管理、跨语言食谱生成研究

基于菜谱数据的食品图像理解研究因数据多样性和复杂性长期受到关注。食物与日常生活紧密相关,对膳食管理等实际应用具有重要意义。近期多模态大语言模型(MLLMs)展现出强大能力,不仅知识广博,且能自然处理语言。尽管主流使用英语,但也能支持包括日语在内的多种语言。这表明MLLMs有望显著提升食品图像理解性能。我们对开源模型LLaVA-1.5和Phi-3 Vision在日式菜谱数据集上进行微调,并与闭源模型GPT-4o进行对比。通过涵盖日本饮食文化的5000个样本评估生成菜谱的内容,包括食材和烹饪步骤。评估显示,微调后的开源模型在食材生成上超越了当前最先进模型GPT-4o,F1分数达到0.531,高于GPT-4o的0.481,准确率更高。同时,在生成烹饪步骤文本方面表现与GPT-4o相当。

原文摘要 · Abstract (English)

Research on food image understanding using recipe data has been a long-standing focus due to the diversity and complexity of the data. Moreover, food is inextricably linked to people's lives, making it a vital research area for practical applications such as dietary management. Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities, not only in their vast knowledge but also in their ability to handle languages naturally. While English is predominantly used, they can also support multiple languages including Japanese. This suggests that MLLMs are expected to significantly improve performance in food image understanding tasks. We fine-tuned open MLLMs LLaVA-1.5 and Phi-3 Vision on a Japanese recipe dataset and benchmarked their performance against the closed model GPT-4o. We then evaluated the content of generated recipes, including ingredients and cooking procedures, using 5,000 evaluation samples that comprehensively cover Japanese food culture. Our evaluation demonstrates that the open models trained on recipe data outperform GPT-4o, the current state-of-the-art model, in ingredient generation. Our model achieved F1 score of 0.531, surpassing GPT-4o's F1 score of 0.481, indicating a higher level of accuracy. Furthermore, our model exhibited comparable performance to GPT-4o in generating cooking procedure text.

多模态模型食谱生成日语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。