arXiv:2601.07581cs.CV2026-01被引 1

构建首个360°多视角食物视频分割数据集,解决新视角下分割失准问题。

BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation

  • 整合55个菜系场景,2.5万帧精细标注,覆盖自由360°相机运动。
  • 记忆增强模型在新视角下保持时间一致性,比FoodMem提升2.63% mAP。
  • 提出时序评估协议,量化分割稳定性,揭示传统方法忽略的失效模式。

食物图像分割对饮食分析至关重要,可实现食物体积与营养成分的精准估算。然而,现有方法受限于多视角数据不足,且对新视角泛化能力差。本文提出BenchSeg,一个全新的多视角食物视频分割数据集与基准测试。BenchSeg整合了来自Nutrition5k、Vegetables & Fruits、MetaFood3D和FoodKit的55个菜品场景,共25,284帧精细标注,捕捉每个菜品在自由360°相机运动下的表现。我们在FoodSeg103上评估20种先进分割模型(如SAM-based、transformer、CNN及大型多模态模型),并在BenchSeg上评估其单独表现及与视频记忆模块结合的效果。定量与定性结果表明,标准图像分割器在新视角下性能显著下降,而记忆增强方法能保持帧间一致性。最佳模型基于SeTR-MLA+XMem2组合,相比先前工作(如FoodMem)mAP提升约2.63%,为饮食分析中的食物分割与追踪提供新见解。除逐帧空间精度外,我们引入专用时序评估协议,通过连续性、闪烁率和IoU漂移等指标显式量化分割稳定性,揭示传统单帧评估难以发现的失败模式。BenchSeg已公开发布,项目页面含数据标注与分割模型:https://amughrabi.github.io/benchseg。

原文摘要 · Abstract (English)

Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view data and poor generalization to new viewpoints. We introduce BenchSeg, a novel multi-view food video segmentation dataset and benchmark. BenchSeg aggregates 55 dish scenes (from Nutrition5k, Vegetables & Fruits, MetaFood3D, and FoodKit) with 25,284 meticulously annotated frames, capturing each dish under free 360° camera motion. We evaluate a diverse set of 20 state-of-the-art segmentation models (e.g., SAM-based, transformer, CNN, and large multimodal) on the existing FoodSeg103 dataset and evaluate them (alone and combined with video-memory modules) on BenchSeg. Quantitative and qualitative results demonstrate that while standard image segmenters degrade sharply under novel viewpoints, memory-augmented methods maintain temporal consistency across frames. Our best model based on a combination of SeTR-MLA+XMem2 outperforms prior work (e.g., improving over FoodMem by ~2.63% mAP), offering new insights into food segmentation and tracking for dietary analysis. In addition to frame-wise spatial accuracy, we introduce a dedicated temporal evaluation protocol that explicitly quantifies segmentation stability over time through continuity, flicker rate, and IoU drift metrics. This allows us to reveal failure modes that remain invisible under standard per-frame evaluations. We release BenchSeg to foster future research. The project page including the dataset annotations and the food segmentation models can be found at https://amughrabi.github.io/benchseg.

食物分割视频分割多视角时序评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。