arXiv:2601.20303cs.CVcs.AI2026-01中稿 · IJCAI

从单张图片估算物体质量,结合几何与材质信息提升准确性

Physically Guided Visual Mass Estimation from a Single RGB Image

  • 用单目深度估计重建3D几何,结合视觉语言模型提取材质语义
  • 在image2mass和ABO-500数据集上优于现有方法,误差显著降低
  • 适合需要物理感知的视觉理解任务,如机器人抓取与场景分析

从视觉输入中估算物体质量具有挑战性,因为质量同时依赖于几何体积和材料相关的密度,而这两者均无法直接从RGB图像中获取。因此,仅凭像素进行质量预测是病态问题,需借助物理上合理的表征来约束可能解的空间。本文提出一种物理结构化的单图像质量估计框架,通过将视觉线索与决定质量的物理因素对齐,解决该模糊性。从单张RGB图像出发,利用单目深度估计恢复以物体为中心的三维几何以估算体积,并通过视觉语言模型提取粗粒度材料语义以指导密度相关推理。这些几何、语义和外观表示通过实例自适应门控机制融合,两个物理引导的潜在因子(体积相关与密度相关)在仅有质量监督下由独立回归头预测。在image2mass和ABO-500数据集上的实验表明,所提方法持续优于现有最先进方法。

原文摘要 · Abstract (English)

Estimating object mass from visual input is challenging because mass depends jointly on geometric volume and material-dependent density, neither of which is directly observable from RGB appearance. Consequently, mass prediction from pixels is ill-posed and therefore benefits from physically meaningful representations to constrain the space of plausible solutions. We propose a physically structured framework for single-image mass estimation that addresses this ambiguity by aligning visual cues with the physical factors governing mass. From a single RGB image, we recover object-centric three-dimensional geometry via monocular depth estimation to inform volume and extract coarse material semantics using a vision-language model to guide density-related reasoning. These geometry, semantic, and appearance representations are fused through an instance-adaptive gating mechanism, and two physically guided latent factors (volume- and density-related) are predicted through separate regression heads under mass-only supervision. Experiments on image2mass and ABO-500 show that the proposed method consistently outperforms state-of-the-art methods.

质量估计3D重建视觉语言模型物理先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。