arXiv:2602.19063cs.CV2026-02

解决3D多模态模型方向感知难题,自动补全位置姿态提升推理准确率。

Direction-aware 3D Large Multimodal Models

  • 自动恢复视角姿态,通过视觉与点云交集匹配定位相机位姿。
  • 点云数据按识别姿态重对齐,使模型更准确理解空间方向。
  • 适配多种3D模型,仅需指令微调,性能显著提升且通用性强。

3D大型多模态模型(3D LMMs)严重依赖自车姿态实现方向性问答与空间推理。然而,现有点云基准数据集虽包含丰富方向性问题,却缺乏对应自车姿态,导致建模本质不成立。本文提出全新严谨范式,通过识别并补充自车姿态,将点云数据按识别姿态重对齐,实现方向感知的3D LMM。设计两个新方法:PoseRecover为全自动姿态恢复流水线,基于RGB-D视频外参,通过物体锥体交集与Z缓冲可见性检测匹配问题与姿态;PoseAlign将点云数据变换至与识别姿态一致,而非注入提示或编码特征。大量实验表明,该方法在多个3D LMM骨干模型(如LL3DA、LL3DA-SONATA、Chat-Scene、3D-LLAVA)上均带来稳定提升,ScanRefer mIoU提高30.0%,Scan2Cap LLM-as-judge准确率提升11.7%。方法简洁、通用、训练高效,仅需指令微调即可建立强基线。

原文摘要 · Abstract (English)

3D large multimodal models (3D LMMs) rely heavily on ego poses for enabling directional question-answering and spatial reasoning. However, most existing point cloud benchmarks contain rich directional queries but lack the corresponding ego poses, making them inherently ill-posed in 3D large multimodal modelling. In this work, we redefine a new and rigorous paradigm that enables direction-aware 3D LMMs by identifying and supplementing ego poses into point cloud benchmarks and transforming the corresponding point cloud data according to the identified ego poses. We enable direction-aware 3D LMMs with two novel designs. The first is PoseRecover, a fully automatic pose recovery pipeline that matches questions with ego poses from RGB-D video extrinsics via object-frustum intersection and visibility check with Z-buffers. The second is PoseAlign that transforms the point cloud data to be aligned with the identified ego poses instead of either injecting ego poses into textual prompts or introducing pose-encoded features in the projection layers. Extensive experiments show that our designs yield consistent improvements across multiple 3D LMM backbones such as LL3DA, LL3DA-SONATA, Chat-Scene, and 3D-LLAVA, improving ScanRefer mIoU by 30.0% and Scan2Cap LLM-as-judge accuracy by 11.7%. In addition, our approach is simple, generic, and training-efficient, requiring only instruction tuning while establishing a strong baseline for direction-aware 3D-LMMs.

3D多模态空间推理姿态恢复点云对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。