用强化学习提升食物视觉语言模型的推理能力
Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

- 采用多任务学习与思维链指令微调,增强模型理解力
- 在8万张食物图像上训练,准确率超越现有方法
- 适合营养分析、饮食建议等实际应用开发者
近期研究探索了视觉-语言模型(VLM)在食物分析中的应用。然而,大多数方法依赖监督微调(SFT),限制了推理与泛化能力。高质量大规模营养标注数据依然稀缺。为此,我们构建了CalorieBench-80K,一个包含8万张带热量标签和饮食建议标注的大规模基准数据集,是首个引入思维链(CoT)标注的食物图像基准。我们提出Food-R1,一种统一的食物视觉语言模型,通过基于思维链的冷启动指令微调,再使用组相对策略优化(GRPO)进行强化微调,以提升推理与性能。在CalorieBench-80K及代表性基准上的实验表明,Food-R1在各类食物相关任务中均持续优于强基线。代码、模型权重与标注数据已开源。
原文摘要 · Abstract (English)
Recent studies have explored Vision-Language Models (VLMs) for food analysis. However, most existing methods rely primarily on supervised fine-tuning (SFT), which often limits reasoning and generalization capabilities. Moreover, high-quality large-scale nutritional annotations remain scarce. To address these issues, we introduce CalorieBench-80K, a large-scale benchmark with curated calorie labels and dietary advice annotations. To the best of our knowledge, it is the first food image benchmark to incorporate Chain-of-Thought (CoT) annotations for calorie reasoning. We also propose Food-R1, a unified food VLM trained in a multi-task learning paradigm to equip the model with broad capabilities. Food-R1 undergoes CoT-based cold-start instruction tuning, followed by reinforcement fine-tuning (RFT) using Group Relative Policy Optimization (GRPO) to improve reasoning and performance. Experiments on CalorieBench-80K and representative benchmarks show that Food-R1 consistently outperforms strong baselines across food-related tasks. The code, model weights, and benchmark annotations are available at the project repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。