arXiv:2607.22355cs.CVcs.AI2026-07中稿 · ECCV

仅凭一张图就能推断物体质量、刚度等物理属性,突破传统多视角依赖。

SiPhy: Single-Image Physical Property Reasoning

论文配图:SiPhy: Single-Image Physical Property Reasoning
图 1 · 摘自论文原文
  • 用视觉与语言模型结合,从单张图提取材料候选并对齐3D特征。
  • 在多个数据集上质量预测误差降低93%,密度估计误差减少35.5%。
  • 适合需要物理理解的机器人、仿真系统,尤其擅长密集物体建模。

从单张图像推断质量、刚度和弹性等物理属性对模拟和具身智能至关重要,但现有方法多依赖多视角重建或基于物理的监督。本文提出SiPhy,一种统一的单图像物理属性推理框架,通过融合3D感知视觉线索、深度信息与基于语言的材料知识。仅需一张RGB图像,SiPhy采样伪体素点,提取CLIP特征,并将其与视觉语言模型(VLM)提出的材料候选进行对齐。基于部件的对比聚合器确保区域一致性,而重量感知优化提升了密集物体的厚度与体积估计。在ABO-500、MVImgNet-100和PhysXNet-100数据集上,SiPhy达到单图像推理新纪录,相比多视角重建方法,质量相对均方根误差(MnRE)降低最高达93%(对比PUGS),密度平均绝对误差(MAE)减少35.5%(对比NeRF2Physics),杨氏模量误差降低23.5%。进一步在真实手-物交互数据集上验证,表明其可作为单视角图像中物理理解的数据标注引擎。

原文摘要 · Abstract (English)

Inferring physical properties such as mass, stiffness, and elasticity from a single image is essential for simulation and embodied AI, yet most existing approaches rely on multi-view reconstruction or physics-based supervision. We introduce SiPhy, a unified framework for single-image physical property reasoning that aligns 3D-aware visual cues, depth with language-based material knowledge. From one RGB image, SiPhy samples pseudo-voxel points, extracts CLIP features, and grounds them to material candidates proposed by a VLM. A part-based contrastive aggregator enforces region consistency, while a heaviness-aware refinement improves thickness and volume estimation for dense objects. Across ABO-500, MVImgNet-100, and PhysXNet-100, SiPhy achieves state-of-the-art single-image performance, surpassing multi-view reconstruction methods by improving mass MnRE by up to 93% (vs. PUGS), reducing density MAE by 35.5% (vs. NeRF2Physics), and lowering Young's modulus error by 23.5%. We further validate SiPhy on real hand-object interaction datasets, demonstrating its potential as a data annotation engine for physical understanding from single-view imagery.

物理属性推理单图像视觉语言模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。