提出高效3D视觉语言模型,提升场景理解与鲁棒性。
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
- 用紧凑特征网格降低3D表示的计算开销。
- 在70万+数据上训练,跨4类场景5种任务达领先性能。
- 引入新后训练方法,显著增强模型对复杂场景的适应力。
构建能理解三维场景的视觉语言模型仍是长期研究目标。尽管已有进展,现有3D VLM仍面临空间推理能力弱和鲁棒性差的问题。我们识别出三大障碍:(1) 场景表示受限于容量与效率的权衡,阻碍可扩展学习;(2) 训练数据缺乏全面性,任务与场景域多样性不足;(3) 模型鲁棒性差,缺乏有效后训练机制。为此,我们提出紧凑特征网格(CFG),在大幅减少令牌开销的同时保持强感知能力。基于CFG,我们构建了LEO-VL,该模型在超过70万条涵盖四个真实室内场景域及五项任务(如描述生成、对话)的3D视觉语言数据上进行训练。为进一步提升鲁棒性,我们提出SceneDPO,一种融合答案与场景对比信号的新后训练目标。LEO-VL在SQA3D、Beacon3D、Scan2Cap等多个3D-VL基准上达到当前最优表现。大量分析揭示了CFG的高效性,并指出任务与场景多样性的重要性、数据质量在规模化中的优先级,以及SceneDPO的优势。
原文摘要 · Abstract (English)
Developing vision-language models (VLMs) capable of understanding 3D scenes has been a longstanding research goal. Despite recent progress, 3D VLMs still struggle with spatial reasoning and robustness. We identify three key obstacles hindering their progress: (1) scene representation is constrained by a capacity-efficiency trade-off, which impedes scalable learning; (2) training data lacks a comprehensive scheme, with limited diversity across tasks and scene domains; and (3) models exhibit robustness deficiencies and lack effective post-training. To address these challenges, we first propose condensed feature grid (CFG), an efficient scene representation that significantly reduces token overhead while preserving strong perceptual capacity. Building on CFG, we introduce LEO-VL, a 3D VLM trained on over 700k 3D vision-language (3D-VL) data spanning four real-world indoor domains and five tasks such as captioning and dialogue. To further improve robustness, we propose SceneDPO, a novel post-training objective that incorporates contrastive signals across both answers and scenes. LEO-VL achieves state-of-the-art performance on various 3D-VL benchmarks, such as SQA3D, Beacon3D, and Scan2Cap. Extensive analyses highlight the efficiency of CFG and provide key insights such as the importance of task and scene diversity, the priority of data quality for effective scaling, and the advantages of SceneDPO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。