用视觉变压器分析毫米波雷达微多普勒谱,实现高精度目标分类。
Temporal Micro-Doppler Spectrogram-based ViT Multiclass Target Classification
- 基于时序微多普勒谱设计跨轴注意力机制,建模多帧数据序列特征。
- 在重叠与遮挡下仍保持分类准确率,优于传统CNN方法。
- 可解释性分析聚焦关键能量区域,适合实时部署场景。
本文提出一种基于时序微多普勒谱的视觉变压器(T-MDS-ViT),用于毫米波调频连续波(FMCW)雷达下的多类目标分类。该模型通过堆叠的方位-速度-距离(RVA)时空张量,利用补丁嵌入和跨轴注意力机制,显式建模微多普勒谱在多帧间的时序特性。T-MDS-ViT在注意力层中引入运动感知约束,以维持目标重叠和部分遮挡下的可分性。此外,我们引入可解释机制,分析注意力层对微多普勒表征中高能量区域的关注,及其对类别特定运动特征的影响。实验表明,所提框架在分类精度上优于现有基于CNN的方法,同时具备更高的数据效率和实时部署能力。
原文摘要 · Abstract (English)
In this paper, we propose a new Temporal MDS-Vision Transformer (T-MDS-ViT) for multiclass target classification using millimeter-wave FMCW radar micro-Doppler spectrograms. Specifically, we design a transformer-based architecture that processes stacked range-velocity-angle (RVA) spatiotemporal tensors via patch embeddings and cross-axis attention mechanisms to explicitly model the sequential nature of MDS data across multiple frames. The T-MDS-ViT exploits mobility-aware constraints in its attention layer correspondences to maintain separability under target overlaps and partial occlusions. Next, we apply an explainable mechanism to examine how the attention layers focus on characteristic high-energy regions of the MDS representations and their effect on class-specific kinematic features. We also demonstrate that our proposed framework is superior to existing CNN-based methods in terms of classification accuracy while achieving better data efficiency and real-time deployability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。