arXiv:2510.11996cs.CV2025-10ICCV被引 4

通过增强提示中的空间坐标,提升模型对物体关系的精准推理能力。

Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning

  • 将边界框坐标嵌入提示,显式引导模型理解物体几何与布局
  • 在四个任务上微调,最终在仓库场景评测中取得73.06分,排名第四
  • 适合需要高精度空间推理的工业视觉任务应用

大规模3D环境(如仓库)中的空间推理对视觉语言系统仍是重大挑战,受限于场景杂乱、遮挡及对精确空间理解的需求。现有模型在该类场景中泛化能力差,主要依赖局部外观且缺乏显式空间定位。本文针对2025 AI City Challenge Track 3的Physical AI Spatial Intelligence Warehouse数据集,提出专用空间推理框架。通过将掩码尺寸以边界框坐标形式直接嵌入输入提示,增强模型对物体几何与布局的理解。在距离估计、物体计数、多选定位和空间关系推理四类任务上,采用特定任务监督进行微调。为提升与评估系统的兼容性,训练集中在GPT响应中添加归一化答案。整体流程最终获得73.0606分,位列公开排行榜第4名。结果表明,结构化提示增强与针对性优化可有效推进真实工业环境中空间推理能力。

原文摘要 · Abstract (English)

Spatial reasoning in large-scale 3D environments such as warehouses remains a significant challenge for vision-language systems due to scene clutter, occlusions, and the need for precise spatial understanding. Existing models often struggle with generalization in such settings, as they rely heavily on local appearance and lack explicit spatial grounding. In this work, we introduce a dedicated spatial reasoning framework for the Physical AI Spatial Intelligence Warehouse dataset introduced in the Track 3 2025 AI City Challenge. Our approach enhances spatial comprehension by embedding mask dimensions in the form of bounding box coordinates directly into the input prompts, enabling the model to reason over object geometry and layout. We fine-tune the framework across four question categories namely: Distance Estimation, Object Counting, Multi-choice Grounding, and Spatial Relation Inference using task-specific supervision. To further improve consistency with the evaluation system, normalized answers are appended to the GPT response within the training set. Our comprehensive pipeline achieves a final score of 73.0606, placing 4th overall on the public leaderboard. These results demonstrate the effectiveness of structured prompt enrichment and targeted optimization in advancing spatial reasoning for real-world industrial environments.

空间推理视觉语言工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。