arXiv:2606.04986cs.CV2026-06

用强化学习提升食物视觉语言模型的推理能力

Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

论文配图:Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning
图 1 · 摘自论文原文
  • 采用多任务学习与思维链指令微调,增强模型理解力
  • 在8万张食物图像上训练,准确率超越现有方法
  • 适合营养分析、饮食建议等实际应用开发者

近期研究探索了视觉-语言模型(VLM)在食物分析中的应用。然而,大多数方法依赖监督微调(SFT),限制了推理与泛化能力。高质量大规模营养标注数据依然稀缺。为此,我们构建了CalorieBench-80K,一个包含8万张带热量标签和饮食建议标注的大规模基准数据集,是首个引入思维链(CoT)标注的食物图像基准。我们提出Food-R1,一种统一的食物视觉语言模型,通过基于思维链的冷启动指令微调,再使用组相对策略优化(GRPO)进行强化微调,以提升推理与性能。在CalorieBench-80K及代表性基准上的实验表明,Food-R1在各类食物相关任务中均持续优于强基线。代码、模型权重与标注数据已开源。

原文摘要 · Abstract (English)

Recent studies have explored Vision-Language Models (VLMs) for food analysis. However, most existing methods rely primarily on supervised fine-tuning (SFT), which often limits reasoning and generalization capabilities. Moreover, high-quality large-scale nutritional annotations remain scarce. To address these issues, we introduce CalorieBench-80K, a large-scale benchmark with curated calorie labels and dietary advice annotations. To the best of our knowledge, it is the first food image benchmark to incorporate Chain-of-Thought (CoT) annotations for calorie reasoning. We also propose Food-R1, a unified food VLM trained in a multi-task learning paradigm to equip the model with broad capabilities. Food-R1 undergoes CoT-based cold-start instruction tuning, followed by reinforcement fine-tuning (RFT) using Group Relative Policy Optimization (GRPO) to improve reasoning and performance. Experiments on CalorieBench-80K and representative benchmarks show that Food-R1 consistently outperforms strong baselines across food-related tasks. The code, model weights, and benchmark annotations are available at the project repository.

食物视觉视觉语言模型强化学习多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。