用几何增强方法提升多模态大模型的食量估计精度
Geometry-Enhanced Portion Estimation for Multimodal LLMs

- 在冻结的多模态大模型基础上添加轻量几何头,结合边界框和密度信息
- 食量估计误差降低33%-41%,优于所有旗舰模型直接输出结果
- 无需深度传感器或微调,适合真实场景饮食评估应用
基于图像的饮食评估有望替代成本高、易偏倚的手动回忆,但食量估计仍是主要障碍。多模态大模型(MLLMs)可在非受控照片中零样本识别多种食物,但在食量估计方面表现薄弱——我们在当前前沿模型(Gemini、GPT、Claude等旗舰版)上测量了这一差距。本文提出一种方法:在冻结的商用MLLM基础上,加入一个精确的食量估计头——基于冻结DINOv2主干网络的小型几何增强网络,采用结构化softmax-所有权体积,输入MLLM输出的食物名称、边界框和密度范围,无需深度传感器,也无需对MLLM进行微调。在三个真实世界基准上进行全开集评估,该方法使每种食物的食量估计误差相对降低33%-41%,优于所有旗舰MLLM的直接估计,并超越各基准原发表的仅图像模型在各自报告指标下的表现。
原文摘要 · Abstract (English)
Image-based dietary assessment promises to replace costly, bias-prone manual recalls, but portion estimation remains a major blocker. Multimodal LLMs (MLLMs) recognize a wide range of foods zero-shot in uncontrolled photos, yet they are weak at portion estimation -- a gap we measure across the current frontier (Gemini, GPT, and Claude flagships alike). We present a method that enhances a frozen, commercial MLLM with an accurate portion head: a small geometry-enhanced network on a frozen DINOv2 backbone with a structured softmax-ownership volume, consuming the MLLM's per-food name, bounding box, and density range -- no depth sensor, no MLLM fine-tuning. Evaluated fully open-vocabulary on three real-world benchmarks, the head cuts per-food portion error by 33-41% relative to the MLLM alone, outperforms every flagship MLLM's direct estimates, and surpasses each benchmark's originally published image-only model at its own reported metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。