用视频直接理解3D场景,无需额外3D数据。
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
- 通过视频输入结合几何先验,实现3D场景感知。
- 在多个3D任务上超越现有方法,尤其在密集描述和视觉定位上表现突出。
- 适合需要轻量化部署的3D视觉理解应用。
近期多模态大模型在二维视觉-语言推理方面取得显著进展,但将其能力扩展至三维场景理解仍面临挑战。现有三维多模态大模型通常依赖外部3D数据输入,限制了可扩展性和泛化能力。为此,我们提出Vid-LLM,一种基于视频输入的3D多模态大模型,无需依赖额外3D数据,具备实际部署可行性。方法中,直接利用几何先验增强场景感知性能,并设计跨任务适配器(CTA)模块,将3D几何线索与视觉-语言表示对齐。为确保几何一致性与完整性,引入度量深度模型,从重建输出中恢复真实尺度几何结构。最后,采用两阶段蒸馏优化策略进行微调,实现快速收敛与训练稳定。在多个基准测试中,实验验证了该方法在3D问答、3D密集描述和3D视觉定位任务上的有效性,展现出卓越的多任务能力。
原文摘要 · Abstract (English)
Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision-Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge. Existing 3D Multimodal Large Language Models (3D-MLLMs) often depend on 3D data inputs, which limits scalability and generalization. To address this limitation, we propose Vid-LLM, a video-based 3D-MLLM that directly processes video inputs without requiring external 3D data, making it practical for real-world deployment. In our method, the geometric prior are directly used to improve the performance of the sceen perception. To integrate the geometric cues into the MLLM compactly, we design a Cross-Task Adapter (CTA) module to align the 3D geometric priors with the vision-language representations. To ensure geometric consistency and integrity, we introduce a Metric Depth Model that recovers real-scale geometry from the reconstruction outputs. Finally, the model is fine-tuned with a two-stage distillation optimization strategy, realizing fast convergence and stabilizes training. Extensive experiments across diverse benchmarks verified the effectiveness of our method on 3D Question Answering, 3D Dense Captioning and 3D Visual Grounding tasks, demonstrating the superior multi-task capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。