用激光雷达轨迹生成时间动态描述,提升自动驾驶视频理解能力
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
- 基于激光雷达轨迹规则提取车道位置与相对运动信息
- 仅用前视摄像头训练模型,跨三数据集提升时间理解能力
- 适合自动驾驶场景中需要精细时序分析的研究者
近年来,视频字幕模型在捕捉时间信息方面取得显著进展,尤其体现在时间注意力机制等架构创新上。然而,现有研究仍缺乏对模型如何捕获和利用时间语义以实现有效时序特征提取的深入理解,特别是在高级驾驶辅助系统(ADAS)背景下的应用。本文提出一种基于激光雷达的自动化字幕生成方法,聚焦交通参与者的时间动态特性。该方法首先通过规则系统从目标轨迹中提取车道位置、相对运动等关键信息,再结合模板生成字幕。实验表明,使用专为捕捉细粒度时间行为设计的模板字幕监督训练SwinBERT模型,仅依赖前视摄像头图像,即可在三个数据集上持续提升时间理解能力。结果明确显示,引入激光雷达字幕监督能显著增强时间感知,有效缓解当前先进模型中存在的视觉/静态偏差问题。
原文摘要 · Abstract (English)
Video captioning models have seen notable advancements in recent years, especially with regard to their ability to capture temporal information. While many research efforts have focused on architectural advancements, such as temporal attention mechanisms, there remains a notable gap in understanding how models capture and utilize temporal semantics for effective temporal feature extraction, especially in the context of Advanced Driver Assistance Systems. We propose an automated LiDAR-based captioning procedure that focuses on the temporal dynamics of traffic participants. Our approach uses a rule-based system to extract essential details such as lane position and relative motion from object tracks, followed by a template-based caption generation. Our findings show that training SwinBERT, a video captioning model, using only front camera images and supervised with our template-based captions, specifically designed to encapsulate fine-grained temporal behavior, leads to improved temporal understanding consistently across three datasets. In conclusion, our results clearly demonstrate that integrating LiDAR-based caption supervision significantly enhances temporal understanding, effectively addressing and reducing the inherent visual/static biases prevalent in current state-of-the-art model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。