arXiv:2605.08663cs.CV2026-05中稿 · the Proceedings of…

用雷达数据识别人手语,准确率超80%,突破信号处理瓶颈。

CAST: Channel-Aware Spatial Transfer Learning with Pseudo-Image Radar for Sign Language Recognition

论文配图:CAST: Channel-Aware Spatial Transfer Learning with Pseudo-Image Radar for Sign Language Recognition
图 1 · 摘自论文原文
  • 通过物理感知设计,将雷达信号转为可识别的运动图谱。
  • 在5折交叉验证中达80.5%准确率,比最优基线提升3.3%。
  • 适合做无摄像头手势识别的科研与医疗应用。

我们提出CAST,一种双流架构,用于解决仅含幅度的60~GHz雷达距离-时间图(RTM)在孤立手语识别中的挑战。该框架结合三种物理感知结构与预训练视觉骨干网络,在临床与字母手势场景下实现纯雷达输入。首先,采用显式分贝转线性变换与加窗快速傅里叶变换,提取节律速度图(CVD),避免对对数压缩信号进行频谱分析时产生的谐波伪影。其次,跨天线空间注意力模块在卷积前对原始天线通道施加注意力,保留接收器间的幅度协方差。第三,非对称交叉注意力机制融合并行的ConvNeXt-Tiny(CVD)与EfficientNetV2-S(RTM)骨干网络表征。大量实验表明,该架构在5折交叉验证下达到80.5%的Top-1准确率,较最优单模型基线(77.2%)提升3.3%。结果表明,物理感知信号表示为受限传感器模态下的纯雷达手语识别提供了有前景的方向。代码已开源:https://github.com/Shakhoyat/CAST-at-SignEval2026。

原文摘要 · Abstract (English)

We propose CAST, a dual-stream architecture that utilizes channel-aware spatial transfer learning for isolated sign language recognition addressing the challenges of magnitude-only 60~GHz radar Range-Time Maps (RTM). The proposed framework combines three physics-aware architectures with pretrained vision backbones, which operate under radar-only constraints across clinical and alphabetical gestures. First, an explicit decibel-to-linear inversion is combined with a windowed fast Fourier transform that extracts Cadence Velocity Diagrams (CVD) while avoiding the harmonic artifacts that arise from the spectral analysis of log-compressed signals. Second, a cross-antenna spatial attention module applies attention to raw antenna channels before the convolution, preserving inter-receiver amplitude covariance. Third, an asymmetric cross-attention mechanism fuses representations from parallel ConvNeXt-Tiny (CVD) and EfficientNetV2-S (RTM) backbones. Extensive experiments reveal that the architecture achieves a Top-1 accuracy of 80.5% under 5-fold cross-validation, establishing a 3.3% improvement over the best single-model baseline (77.2%). The findings suggest that physics-aware signal representations form a promising direction for radar-only sign language recognition under constrained sensor modalities. The source code is available at: https://github.com/Shakhoyat/CAST-at-SignEval2026.

手语识别雷达感知多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。