arXiv:2604.06352cs.CVcs.AI2026-04被引 1

用前后餐食图像和语言提示,精准估算吃掉的食物重量。

DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images

  • 用语言提示定位食物,直接从单张图像估重
  • 通过双阶段训练预测图像间重量差,实现精准摄入量估计
  • 无需深度图或分割,适合普通用户日常使用

准确的膳食评估对精准营养至关重要,但现有基于图像的方法多依赖单张进食前图像,仅能提供粗略的餐食级估计,无法确定实际摄入内容,且常需深度感知、多视角图像或显式分割等限制性输入。本文提出一种简单的视觉-语言框架,利用成对的进食前后图像进行食物级营养分析。该方法不依赖刚性分割掩码,而是借助自然语言提示定位特定食物并直接从单张RGB图像中估算其重量。进一步通过两阶段训练策略预测成对图像间的重量差异,实现食物消耗量估计。我们在三个公开数据集上评估该方法,结果表明其在各项指标上均优于现有方法,为前后餐食图像分析建立了强基准。

原文摘要 · Abstract (English)

Accurate dietary assessment is critical for precision nutrition, yet most image-based methods rely on a single pre-consumption image and provide only coarse, meal-level estimates. These approaches cannot determine what was actually consumed and often require restrictive inputs such as depth sensing, multi-view imagery, or explicit segmentation. In this paper, we propose a simple vision-language framework for food-item-level nutritional analysis using paired before-and-after eating images. Instead of relying on rigid segmentation masks, our method leverages natural language prompts to localize specific food items and estimate their weight directly from a single RGB image. We further estimate food consumption by predicting weight differences between paired images using a two-stage training strategy. We evaluate our method on three publicly available datasets and demonstrate consistent improvements over existing approaches, establishing a strong baseline for before-and-after dietary image analysis.

膳食评估视觉语言图像分析营养分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。