arXiv:2504.09862cs.LG2025-04AAAI被引 13

用大模型理解毫米波雷达点云,实现隐私保护下的动作识别

RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-Wave Point Cloud Sequence

  • 通过可变形身体模板和轨迹建模,将雷达点云转为语义标记
  • 在真实与合成数据上均达领先效果,恶劣环境下仍稳定准确
  • 适合关注隐私感知、多模态融合的智能感知研究者

毫米波雷达提供了一种隐私友好且环境鲁棒的感知方式,可在低光、遮挡、雨雪或烟雾等复杂条件下进行人体动作分析。然而其稀疏点云给语义理解带来挑战。本文提出RadarLLM,首个利用大语言模型(LLM)理解雷达信号中人体动作的框架。该框架包含两项关键创新:(1) 基于自研聚合VQ-VAE架构的运动引导雷达分词器,融合可变形身体模板与掩码轨迹建模,将时空雷达序列转换为紧凑语义标记;(2) 雷达感知语言模型,在共享嵌入空间中建立雷达与文本的跨模态对齐。为缓解雷达-文本配对数据稀缺问题,采用物理感知合成流程从运动-文本数据集中生成真实感雷达-文本数据集。在合成与真实世界基准上的大量实验表明,RadarLLM达到当前最优性能,实现了在隐私与可视性受限条件下的鲁棒且可解释的动作理解,即使在恶劣环境中亦表现良好。

原文摘要 · Abstract (English)

Millimeter-wave radar offers a privacy-preserving and environment-robust alternative to vision-based sensing, enabling human motion analysis in challenging conditions such as low light, occlusions, rain, or smoke. However, its sparse point clouds pose significant challenges for semantic understanding. We present RadarLLM, the first framework that leverages large language models (LLMs) for human motion understanding from radar signals. RadarLLM introduces two key innovations: (1) a motion-guided radar tokenizer based on our Aggregate VQ-VAE architecture, integrating deformable body templates and masked trajectory modeling to convert spatial-temporal radar sequences into compact semantic tokens; and (2) a radar-aware language model that establishes cross-modal alignment between radar and text in a shared embedding space. To overcome the scarcity of paired radar-text data, we generate a realistic radar-text dataset from motion-text datasets with a physics-aware synthesis pipeline. Extensive experiments on both synthetic and real-world benchmarks show that RadarLLM achieves state-of-the-art performance, enabling robust and interpretable motion understanding under privacy and visibility constraints, even in adverse environments. This paper has been accepted for presentation at AAAI 2026. This is an extended version with supplementary materials.

雷达感知大模型动作识别多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。