动态剪枝空间令牌,让3D多模态模型更快更省算力
AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning
- 通过分析注意力模式动态裁剪冗余空间令牌
- 推理速度提升21%,计算量减少63%且精度不变
- 适合追求高效3D多模态模型的开发者
大型多模态模型(LMMs)在3D场景理解中表现突出,但现有3D LMMs使用数千个空间令牌进行多模态推理,导致计算开销大、信息冗余。与处理单张图像的2D VLMs不同,3D LMMs因空间令牌与视觉令牌机制异质性存在固有冗余。为此,我们提出AdaToken-3D,一种自适应空间令牌优化框架,通过空间贡献分析动态剪枝冗余令牌。该方法利用注意力模式挖掘量化令牌级信息流,自动适配不同3D LMM架构。在LLaVA-3D(7B参数3D-LMM)上的实验表明,该方法实现21%的推理加速和63%的FLOPs降低,同时保持原有任务准确率。此外,本工作通过定量令牌交互分析系统揭示了多模态空间信息流中的冗余规律:超过60%的空间令牌对最终预测贡献不足5%,为高效3D多模态学习提供了理论基础。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have become a pivotal research focus in deep learning, demonstrating remarkable capabilities in 3D scene understanding. However, current 3D LMMs employing thousands of spatial tokens for multimodal reasoning suffer from critical inefficiencies: excessive computational overhead and redundant information flows. Unlike 2D VLMs processing single images, 3D LMMs exhibit inherent architectural redundancy due to the heterogeneous mechanisms between spatial tokens and visual tokens. To address this challenge, we propose AdaToken-3D, an adaptive spatial token optimization framework that dynamically prunes redundant tokens through spatial contribution analysis. Our method automatically tailors pruning strategies to different 3D LMM architectures by quantifying token-level information flows via attention pattern mining. Extensive experiments on LLaVA-3D (a 7B parameter 3D-LMM) demonstrate that AdaToken-3D achieves 21\% faster inference speed and 63\% FLOPs reduction while maintaining original task accuracy. Beyond efficiency gains, this work systematically investigates redundancy patterns in multimodal spatial information flows through quantitative token interaction analysis. Our findings reveal that over 60\% of spatial tokens contribute minimally ($<$5\%) to the final predictions, establishing theoretical foundations for efficient 3D multimodal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。