融合雷达红外等多模态信号,让智能设备更懂人的行为。
HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning
- 用雷达、红外等非视觉传感器融合语言理解,提升环境适应性。
- 在新构建数据集上,语言引导的人体感知准确率提升30%。
- 适合做智能家居、人机交互的多模态系统研究者参考。
在智能家居中运行的具身智能体需通过多样感官输入理解人类行为,并以自然语言交流。尽管视觉-语言模型(VLMs)实现了出色的语言引导感知,但其依赖视觉数据,在遮挡、光照差或隐私受限场景下鲁棒性不足。本文提出HoloLLM,一种整合了激光雷达(LiDAR)、红外、毫米波雷达(mmWave radar)和WiFi等罕见但强大的传感模态的多模态大语言模型(MLLM),实现跨异构环境的无缝人体感知与推理。针对两类关键挑战:(1)稀有传感器对齐模态-文本数据稀缺;(2)物理信号表示异质性强。我们设计通用模态注入投影器(UMIP),通过粗到细的交叉注意力机制,将定制编码器提取的细粒度文本对齐特征注入预对齐的模态嵌入,无需显著增加对齐开销。此外,提出人-视觉语言模型协同数据清洗流程,生成传感数据集的配对文本注释。在两个新构建基准上的大量实验表明,HoloLLM显著优于现有MLLM,语言引导的人体感知准确率最高提升30%。本工作为现实世界中语言驱动的多感官具身智能建立了新基础。
原文摘要 · Abstract (English)
Embodied agents operating in smart homes must understand human behavior through diverse sensory inputs and communicate via natural language. While Vision-Language Models (VLMs) have enabled impressive language-grounded perception, their reliance on visual data limits robustness in real-world scenarios with occlusions, poor lighting, or privacy constraints. In this paper, we introduce HoloLLM, a Multimodal Large Language Model (MLLM) that integrates uncommon but powerful sensing modalities, such as LiDAR, infrared, mmWave radar, and WiFi, to enable seamless human perception and reasoning across heterogeneous environments. We address two key challenges: (1) the scarcity of aligned modality-text data for rare sensors, and (2) the heterogeneity of their physical signal representations. To overcome these, we design a Universal Modality-Injection Projector (UMIP) that enhances pre-aligned modality embeddings with fine-grained, text-aligned features from tailored encoders via coarse-to-fine cross-attention without introducing significant alignment overhead. We further introduce a human-VLM collaborative data curation pipeline to generate paired textual annotations for sensing datasets. Extensive experiments on two newly constructed benchmarks show that HoloLLM significantly outperforms existing MLLMs, improving language-grounded human sensing accuracy by up to 30%. This work establishes a new foundation for real-world, language-informed multisensory embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。