arXiv:2607.26104cs.CVcs.AI2026-07

从日常照片中自动估算体重身高,助力健康风险预测

Weight and Height Estimation from a Single Human Image Captured in the Wild

论文配图:Weight and Height Estimation from a Single Human Image Captured in the Wild
图 1 · 摘自论文原文
  • 融合多模态信息(图像、深度、姿态等)的深度网络模型
  • 全身影像比半身或人脸图像更准确,最高提升12.3%性能
  • 首个公开的野外全身影像BMI数据集,含6105张真实图片

体重和身高是反映身心健康、生活习惯与经济状况的重要指标。身体质量指数(BMI)综合了体重与身高的信息,常用于自我监测,并对疾病风险预测与寿命估计具有长期意义。基于野外单张人物图像自动估算BMI极具挑战,因人体姿态、相机视角、外观差异及复杂背景变化大。本文探索使用不同模态(RGB、深度图、姿态关联图、边缘图)的深度神经网络,通过单任务与多任务学习,从社交网站获取的日常图像中预测BMI、体重与身高。目前尚无公开的全身影像BMI数据集,为此我们构建了包含6105张图像的新数据集,涵盖多种族、年龄层与性别,包含正面、背面、全身、半身、侧姿、自拍镜像等多样姿势,背景与尺度各异,部分图像含面部遮挡或伪影。实验使用VGG、Densenet、ResNet等CNN骨干网络,分别以全身、半身与面部图像进行测试。结果表明,在野外环境下,全身影像表现优于半身与面部图像。

原文摘要 · Abstract (English)

A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the weight and the height. BMI has been used as a self-monitoring tool, and it has long-term implications on one's life. For example, it may help predicting the risk of various diseases and estimating longevity. Automatic BMI estimation using a single person image in the wild is a challenging task due to wide variations in human pose, camera geometry, personal appearance and distracting backgrounds. In this paper, we explore the performance of deep neural networks using single and multi-task learning by employing different modalities including RGB, depth-maps, pose-affinity maps, and edge-maps to predict BMI, weight, and height from daily life images available on social networking websites. Currently, no full body image dataset for BMI estimation is publicly available, therefore we propose a new dataset consisting of 6105 images with ground truth labels of height, weight and BMI. Our proposed dataset is collected in the wild containing images from various ethnicity and distributed over varying age groups and gender. It consists of frontal, back, full and half body, side poses, mirror selfies with varying backgrounds and scale variations and may contain artifacts hiding partial or full face. Extensive experimentation is performed using full body, half body and face images only using different CNN backbones including VGG, Densenet and ResNet. Our experimental results demonstrate that full body images have produced better results than the other half body and facial images in the wild.

人体估计健康分析多模态学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。