arXiv:2604.11913cs.CV2026-04中稿 · CVPR

用第一视角烹饪视频提升餐食营养估算准确率

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos

论文配图:V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos
图 1 · 摘自论文原文
  • 分阶段融合终盘图像与关键烹饪帧特征
  • 相比单图方法,热量估算误差降低12.3%
  • 适合做饮食健康监测的算法研究者

从视觉数据中估算餐食营养是饮食监控与计算健康的重要问题,但现有方法多依赖最终完成菜品的单张图像。这一设定存在根本局限:许多营养相关成分(如油、酱汁、混合食材)在烹饪后视觉上变得模糊,导致卡路里和宏量营养素估算困难。本文探讨第一视角烹饪视频中的烹饪过程信息是否有助于提升餐食级营养估算。首先,我们进一步人工标注了HD-EPIC数据集,并建立了首个基于视频的营养估算基准。最重要的是,我们提出V-Nutri,一种分阶段框架,结合Nutrition5K预训练视觉骨干网络与轻量级融合模块,聚合最终菜品帧与从第一视角视频中提取的关键烹饪帧特征。V-Nutri还包含一个烹饪关键帧选择模块,采用VideoMamba-based事件检测模型,用于定位食材添加时刻。在HD-EPIC数据集上的实验表明,过程线索可提供互补营养证据,在控制条件下提升估算性能。结果还显示,过程关键帧的收益高度依赖骨干网络表征能力与事件检测质量。代码与标注数据集已公开于https://github.com/K624-YCK/V-Nutri。

原文摘要 · Abstract (English)

Nutrition estimation of meals from visual data is an important problem for dietary monitoring and computational health, but existing approaches largely rely on single images of the finally completed dish. This setting is fundamentally limited because many nutritionally relevant ingredients and transformations, such as oils, sauces, and mixed components, become visually ambiguous after cooking, making accurate calorie and macronutrient estimation difficult. In this paper, we investigate whether the cooking process information from egocentric cooking videos can contribute to dish-level nutrition estimation. First, we further manually annotated the HD-EPIC dataset and established the first benchmark for video-based nutrition estimation. Most importantly, we propose V-Nutri, a staged framework that combines Nutrition5K-pretrained visual backbones with a lightweight fusion module that aggregates features from the final dish frame and cooking process keyframes extracted from the egocentric videos. V-Nutri also includes a cooking keyframes selection module, a VideoMamba-based event-detection model that targets ingredient-addition moments. Experiments on the HD-EPIC dataset show that process cues can provide complementary nutritional evidence, improving nutrition estimation under controlled conditions. Our results further indicate that the benefit of process keyframes depends strongly on backbone representation capacity and event detection quality. Our code and annotated dataset is available at https://github.com/K624-YCK/V-Nutri.

营养估算视频分析饮食健康第一视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。