用单张图实现多食物体积估算,突破尺度模糊难题。
Implicit-Scale 3D Reconstruction for Multi-Food Volume Estimation from Monocular Images
- 将食物估重转化为隐式尺度的3D重建问题
- 顶级方法体积估计MAPE达0.21,几何误差仅5.7 L1 Chamfer距离
- 适合研究视觉几何推理与饮食分析的学者
我们提出一种基于单目图像的隐式尺度3D重建基准数据集,旨在推动真实就餐场景下的基于几何的食物份量评估。现有饮食评估方法多依赖单图分析或外观推断,包括近期的视觉-语言模型,缺乏显式几何推理且易受尺度模糊影响。本基准将食物份量估计重构为单目观测下的隐式尺度3D重建问题。为贴近真实环境,移除了显式物理参照物和度量标注,转而提供盘子、餐具等上下文物体,要求算法从隐式线索和先验知识中推断尺度。数据集强调多食物场景,包含多样几何形态、频繁遮挡及复杂空间布局。该基准被采纳为MetaFood 2025研讨会挑战任务,多个团队提交基于重建的解决方案。实验表明,尽管强视觉-语言基线表现良好,但基于几何的重建方法在精度和鲁棒性上均更优,最优方法体积估计达到0.21 MAPE,几何准确度为5.7 L1 Chamfer Distance。
原文摘要 · Abstract (English)
We present Implicit-Scale 3D Reconstruction from Monocular Multi-Food Images, a benchmark dataset designed to advance geometry-based food portion estimation in realistic dining scenarios. Existing dietary assessment methods largely rely on single-image analysis or appearance-based inference, including recent vision-language models, which lack explicit geometric reasoning and are sensitive to scale ambiguity. This benchmark reframes food portion estimation as an implicit-scale 3D reconstruction problem under monocular observations. To reflect real-world conditions, explicit physical references and metric annotations are removed; instead, contextual objects such as plates and utensils are provided, requiring algorithms to infer scale from implicit cues and prior knowledge. The dataset emphasizes multi-food scenes with diverse object geometries, frequent occlusions, and complex spatial arrangements. The benchmark was adopted as a challenge at the MetaFood 2025 Workshop, where multiple teams proposed reconstruction-based solutions. Experimental results show that while strong vision--language baselines achieve competitive performance, geometry-based reconstruction methods provide both improved accuracy and greater robustness, with the top-performing approach achieving 0.21 MAPE in volume estimation and 5.7 L1 Chamfer Distance in geometric accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。