用语言提示+立体视觉,提升物体体积估计精度
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
- 融合立体视觉的隐式3D信息与文本显式先验
- 在多个公开数据集上显著优于纯视觉方法
- 适合需要上下文感知的智能测量场景
从视觉数据中准确估算物体体积是计算机视觉中的长期挑战,具有机器人、物流和智慧健康等重要应用。现有方法常依赖复杂的三维重建流程,或难以应对单视角图像固有的模糊性。为此,我们提出一种新方法,将立体视觉的隐式3D线索与自然语言文本中的显式先验知识相融合。该方法从一对立体图像和包含物体类别及近似体积的描述性文本提示中提取深度特征,通过一个简单而有效的投影层整合为统一的多模态表示,用于回归预测。我们在多个公开数据集上进行了大量实验,结果表明,该文本引导方法显著优于仅依赖视觉的基线模型。研究发现,即使使用简单的文本先验也能有效指导体积估计任务,为构建更具上下文感知能力的视觉测量系统开辟了新路径。
原文摘要 · Abstract (English)
Accurate volume estimation of objects from visual data is a long-standing challenge in computer vision with significant applications in robotics, logistics, and smart health. Existing methods often rely on complex 3D reconstruction pipelines or struggle with the ambiguity inherent in single-view images. To address these limitations, we introduce a new method that fuses implicit 3D cues from stereo vision with explicit prior knowledge from natural language text. Our approach extracts deep features from a stereo image pair and a descriptive text prompt that contains the object's class and an approximate volume, then integrates them using a simple yet effective projection layer into a unified, multi-modal representation for regression. We conduct extensive experiments on public datasets demonstrating that our text-guided approach significantly outperforms vision-only baselines. Our findings show that leveraging even simple textual priors can effectively guide the volume estimation task, paving the way for more context-aware visual measurement systems. Code: https://gitlab.com/viper-purdue/stereo-typical-estimator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。