arXiv:2608.04127cs.CV2026-08

用人体姿态引导雷达模型,实现无接触行为理解。

Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding

论文配图:Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding
图 1 · 摘自论文原文
  • 以3D姿态为训练监督,让雷达学人体结构与运动
  • 在17.9小时真实数据上表现优于现有方法
  • 适合智能交互、隐私保护场景的开发人员

大型语言模型代理需要感知物理环境中的行为。毫米波(mmWave)雷达提供一种隐私友好且非接触的传感方式,但雷达观测难以与语言对齐。现有雷达-语言方法常依赖合成数据或缺乏对人体结构与运动的显式监督。我们提出 mmMind,一种使用同步3D姿态作为仅训练监督的雷达-语言模型。一个时空雷达编码器经过预训练以捕捉身体构型与运动动态,之后移除姿态头,使推理仅依赖雷达。学习到的雷达表征随后与大语言模型对齐,用于行为描述生成与时空问答。我们还引入 mmMind-Bench,一个包含23名参与者在7个室内环境中的17.9小时真实记录的毫米波-语言基准。在描述生成、问答和未见动作泛化任务上的实验表明,mmMind持续优于现有雷达-语言基线,消融实验确认了姿态引导预训练的重要性。

原文摘要 · Abstract (English)

Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and contactless sensing modality, but radar observations are difficult to align with language. Existing radar-language methods often rely on synthetic data or lack explicit supervision for human body structure and motion. We present mmMind, a radar-language model that uses synchronized 3D pose as training-only supervision. A spatio-temporal radar encoder is pretrained to capture body configuration and motion dynamics, after which the pose head is removed so that inference requires radar alone. The learned radar representations are then aligned with an LLM for behavior captioning and spatio-temporal question answering. We also introduce mmMind-Bench, a real-world mmWave-language benchmark containing 17.9 hours of recordings from 23 participants across seven indoor environments. Experiments on captioning, question answering, and unseen-action generalization show that mmMind consistently outperforms existing radar-language baselines, while ablations confirm the importance of pose-guided pretraining.

雷达感知行为理解多模态隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。