用视觉提示提升视觉语言模型对物体质量的推理能力
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
- 通过检测、尺度和截面图增强输入图像,引导模型理解物体大小与结构
- 在真实场景数据上,视觉提示使质量估计准确率显著提升
- 适合机器人抓取与物理交互任务的研究者使用
视觉语言模型(VLMs)在机器人感知与操作中应用日益广泛,但其对操作所需物理属性的推断能力仍有限。特别是物体质量估计对于确定合适的抓握力度和保障安全交互至关重要。然而,现有VLMs缺乏可靠的物理量推理能力,且多数基准未在真实感知条件下评估物理量估计。本文提出PhysQuantAgent框架,结合新构建的VisPhysQuant基准数据集,用于真实物体质量估计。VisPhysQuant包含多视角RGB-D视频及精确质量标注。为提升估计精度,我们引入三种视觉提示方法:目标检测、尺度估计与截面图像生成,帮助模型理解目标物体的尺寸与内部结构。实验表明,视觉提示显著提升了真实数据上的质量估计准确率,证明将空间推理与VLM知识融合对物理推断的有效性。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly applied to robotic perception and manipulation, yet their ability to infer physical properties required for manipulation remains limited. In particular, estimating the mass of real-world objects is essential for determining appropriate grasp force and ensuring safe interaction. However, current VLMs lack reliable mass reasoning capabilities, and most existing benchmarks do not explicitly evaluate physical quantity estimation under realistic sensing conditions. In this work, we propose PhysQuantAgent, a framework for real-world object mass estimation using VLMs, together with VisPhysQuant, a new benchmark dataset for evaluation. VisPhysQuant consists of RGB-D videos of real objects captured from multiple viewpoints, annotated with precise mass measurements. To improve estimation accuracy, we introduce three visual prompting methods that enhance the input image with object detection, scale estimation, and cross-sectional image generation to help the model comprehend the size and internal structure of the target object. Experiments show that visual prompting significantly improves mass estimation accuracy on real-world data, suggesting the efficacy of integrating spatial reasoning with VLM knowledge for physical inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。