融合视觉与WiFi信号,更准预测物体内部材料属性。
Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties

- 用材料槽模型捕捉物体分块一致的材质结构,避免逐体素独立估计。
- 结合电磁仿真生成的WiFi特征,提升对相似外观不同材质的区分能力。
- 在6个基准中4项体素级指标领先,适合需要物理属性精确建模的任务。
仅靠视觉难以准确估计物体的体积机械属性(如杨氏模量、泊松比、密度),因为外观相似的物体可能具有完全不同材质和物理行为。现有方法独立预测各体素属性,忽略真实物体的分块常数材质结构,导致同质区域估计噪声大且不一致,且无法有效解决视觉模糊问题。本文提出ViWi(Vision Meets WiFi)框架,以对象为中心,用一组紧凑的材料槽聚合具有相同材质身份的体素证据,并生成一致的槽级属性预测。为补充视觉信息,ViWi引入通过介电常数与电导率进行WiFi频段电磁仿真的紧凑射频描述符,为材料槽提供全局组成线索,而视觉特征则保留体素级空间定位。在体积机械属性与质量估计基准上,ViWi在六个体素级指标中的四项超越现有最佳方法;其纯视觉变体也全面提升了所有质量估计指标。结果表明,结合对象中心材质结构与互补射频证据,可实现超越纯视觉的更精准、更物理一致的体积属性估计。
原文摘要 · Abstract (English)
Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from vision alone, as visually similar objects may have substantially different material compositions and physical behavior. Existing approaches predict these properties independently across voxels, overlooking the piecewise-constant material structure of real objects and producing noisy or inconsistent estimates for voxels that share the same material, while lacking an explicit mechanism to resolve visual ambiguity. We introduce ViWi (Vision Meets WiFi), an object-centric framework for volumetric mechanical-property estimation. ViWi represents each object using a compact set of material slots that aggregate evidence from voxels with a shared material identity and produce coherent slot-level property predictions. To complement visual appearance, ViWi incorporates a compact RF descriptor generated through WiFi-band electromagnetic simulation using permittivity and conductivity. The RF descriptor conditions the material slots with global composition cues that may be unavailable from images, while visual features preserve voxel-level spatial localization. Across volumetric mechanical-property and mass-estimation benchmarks, ViWi improves over the prior state of the art on four of six per-voxel metrics, while its vision-only variant improves all mass-estimation metrics. These results demonstrate that combining object-centric material structure with complementary RF evidence enables more accurate and physically coherent volumetric property estimation beyond what is possible from visual appearance alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。