用关键帧提示让大模型零样本推理3D空间关系,无需专门训练
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
- 通过视觉相似度等指标选关键帧,结合相机位姿抽象空间结构
- 在ScanQA和SQA3D上达到当前最优零样本性能
- 适合想快速部署3D空间推理的开发者,无需3D数据或微调
本研究提出SpatialPrompting框架,利用现成多模态大模型的涌现推理能力,在三维环境实现零样本空间推理。不同于依赖点云或体素等专用3D输入并需昂贵微调的方法,该框架采用关键帧驱动的提示生成策略,基于视觉-语言相似性、马氏距离、视场角和图像清晰度等指标,从图像序列中选取多样且信息丰富的关键帧,并融合对应相机位姿数据,有效抽象空间关系并推断复杂3D结构。该方法不仅建立了一种利用直观视觉与位置线索的灵活空间推理新范式,还在ScanQA和SQA3D等多个基准数据集上实现了领先于现有方法的零样本性能。所提方法彻底避免了专用3D输入与微调需求,为传统方法提供了更简洁、可扩展的替代方案。
原文摘要 · Abstract (English)
This study introduces SpatialPrompting, a novel framework that harnesses the emergent reasoning capabilities of off-the-shelf multimodal large language models to achieve zero-shot spatial reasoning in three-dimensional (3D) environments. Unlike existing methods that rely on expensive 3D-specific fine-tuning with specialized 3D inputs such as point clouds or voxel-based features, SpatialPrompting employs a keyframe-driven prompt generation strategy. This framework uses metrics such as vision-language similarity, Mahalanobis distance, field of view, and image sharpness to select a diverse and informative set of keyframes from image sequences and then integrates them with corresponding camera pose data to effectively abstract spatial relationships and infer complex 3D structures. The proposed framework not only establishes a new paradigm for flexible spatial reasoning that utilizes intuitive visual and positional cues but also achieves state-of-the-art zero-shot performance on benchmark datasets, such as ScanQA and SQA3D, across several metrics. The proposed method effectively eliminates the need for specialized 3D inputs and fine-tuning, offering a simpler and more scalable alternative to conventional approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。