让自动驾驶视觉问答更精准高效,动态选关键视觉信息
LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
- 按问题需求动态筛选关键视觉片段,减少计算量
- 保持高分辨率细节感知,同时高效处理时间序列信息
- 适合需要实时、精细视觉理解的自动驾驶系统
近期视觉语言模型(VLM)在自动驾驶视觉问答(VQA)中取得进展,支持人车自然交互。然而现有方法多基于静态图像或视频,依赖降采样降低计算开销,导致关键细节丢失,难以有效融合时空信息,影响细粒度感知与时间连贯性,制约决策效果。为此,我们提出LaVida Drive,一种新型高效的自动驾驶VQA框架。该框架在保持高分辨率输入以实现精细视觉感知的同时,整合时序数据。通过保留高分辨率空间输入捕捉细节,使用低分辨率输入进行时序分析以聚焦运动特征,提升计算效率。其核心包含两个模块:查询感知的视觉标记选择模块(Query-aware Token Selection),根据问题语义动态筛选最相关视觉标记,减少高分辨率输入的标记数量;以及时空标记恢复与增强模块(Spatial-Temporal Token Recovery and Enhancement),保障空间与时间信息间平滑连贯的交互,维持帧间上下文连续性。在多个自动驾驶问答基准测试中,实验表明LaVida Drive显著减少视觉标记数量,提升效率,并整体性能更优。
原文摘要 · Abstract (English)
Recent advancements in Visual Language Models (VLMs) have made them crucial for visual question answering (VQA) in autonomous driving, enabling natural human-vehicle interactions. However, existing methods often struggle in dynamic driving environments, as they usually focus on static images or videos and rely on downsampling to manage computational costs. This results in the loss of critical details and the difficulty in effectively integrating spatial and temporal information, undermining fine-grained perception and temporal coherence essential for effective decision-making. To tackle these challenges, we introduce LaVida Drive, a novel and efficient VQA framework for autonomous driving. LaVida Drive seamlessly integrates temporal data while maintaining high-resolution inputs for detailed visual perception. It optimizes spatial processing by retaining high-resolution data for intricate details and using lower-resolution inputs for temporal analysis to focus on motion-related features, thereby boosting computational efficiency. The core of LaVida Drive consists of two modules: the \textit{Query-aware Token Selection} module and the \textit{Spatial-Temporal Token Recovery and Enhancement} module. The former dynamically selects the most relevant visual tokens based on semantic alignment with the input query, reducing the token count from high-resolution spatial input. The latter ensures smooth and coherent interactions between spatial and temporal information, preserving contextual continuity across frames. Extensive experiments on various autonomous driving question-answering benchmarks show that LaVida Drive significantly reduces visual tokens, enhances efficiency, and improves overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。