首个真实世界食谱生成多模态基准,支持图文视频生成
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
- 构建了步对齐的食谱多模态数据集,涵盖2.6万食谱与近20万图像
- 提出食材保真度与交互建模评估指标,验证生成质量
- 适合研究食谱生成、跨模态内容创作的研究者使用
食谱图像生成是食品计算中的关键挑战,应用于烹饪教育和多模态食谱助手。然而现有数据集缺乏食谱目标、分步说明与视觉内容之间的细粒度对齐。我们提出RecipeGen,首个大规模真实世界基准,支持基于食谱的文本到图像(T2I)、图像到视频(I2V)和文本到视频(T2V)生成。RecipeGen包含26,453个食谱、196,724张图像和4,491个视频,覆盖多样化食材、烹饪流程、风格与菜品类别。我们进一步提出领域特定评估指标,用于衡量食材保真度与交互建模能力,对代表性T2I、I2V和T2V模型进行基准测试,并为未来食谱生成模型提供洞见。
原文摘要 · Abstract (English)
Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,453 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available now.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。