让大模型理解毫米波雷达数据,首次建立评测基准
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding

- 将雷达点云转为自然语言,让现成大模型直接问答
- 涵盖六种场景五类任务,统一异构雷达数据集
- 展示大模型在遮挡下仍具推理能力,适合智能感知研究
大语言模型(LLM)展现出强大的推理与生成能力,推动其作为通用感知推理引擎的应用。尽管视觉-语言模型已尝试将推理能力融入视觉感知,但毫米波(mmWave)模态——具有低光与遮挡下的独特优势——与大模型的结合仍鲜有探索。主要瓶颈在于雷达-语言配对数据稀缺、跨数据集差异大,以及缺乏基础毫米波编码器。本文通过最小化文本化接口,将每个毫米波点云序列化为简洁自然语言,使现成大模型可在问答(QA)场景中运行。基于此,提出 mmWave-QA,首个面向语言条件毫米波人体感知的基准。该基准整合异构公开毫米波数据集,通过校准感知预处理与全局分类体系对齐实现标准化,并提供自然语言问答。覆盖六种场景与五类任务,支持在不同硬件与实验条件下标准化评估,为毫米波-大模型融合研究奠定基础。进一步在 mmWave-QA 上评估与分析大模型,揭示其在雷达感知中的零样本推理潜力及在视觉退化下的鲁棒性。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for perception. While modern approaches such as vision-language models (VLMs) have attempted to incorporate reasoning capabilities into visual sensing, the integration of LLMs with the millimeter-wave (mmWave) modality-despite its unique advantages under low light and occlusion-remains largely unexplored. The principal bottlenecks stem from the scarcity of radar language pairs, severe cross-dataset heterogeneity, and the absence of a foundational mmWave encoder. We address this gap through a minimal textualization interface that serializes each mmWave point cloud into concise natural language, allowing off-the-shelf LLMs to operate in a question answering (QA) setting. Building on this, we present mmWave-QA, the first benchmark for language-conditioned mmWave human perception. mmWave-QA aggregates heterogeneous public mmWave datasets and harmonizes them via calibration-aware preprocessing and global taxonomy alignment, while providing natural language QA. Spanning six scenarios and five QA tasks, the benchmark enables standardized evaluation across diverse mmWave hardware and experimental conditions, establishing a foundation for scalable research on mmWave-LLM integration. We further evaluate and analyze LLMs on our mmWave-QA, highlighting their zero-shot reasoning potential for radar perception, as well as their robustness under visual degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。