arXiv:2410.10167cs.CVeess.SP2024-10ICLR被引 21

X-Fi让多种传感器任意组合使用,无需重新训练。

X-Fi: A Modality-Invariant Foundation Model for Multimodal Human Sensing

  • 用Transformer和新融合机制实现跨模态独立或组合使用
  • 在六种模态上对姿态与动作识别达当前最优性能
  • 适合需要灵活适配传感器的智能安防与机器人场景

人类感知通过多种传感器与深度学习技术精准捕捉并解析人体信息,在公共安全与机器人等领域产生深远影响。然而,现有方法多依赖相机、LiDAR等固定模态,且多模态融合通常针对特定组合设计,新增或移除模态需大量重训。本文提出通用模态无关基础模型X-Fi,利用Transformer结构支持可变输入规模,并引入创新的“X-fusion”机制,在多模态融合中保留各模态特有特征,提升适应性并促进互补特征学习。在MM-Fi与XRF55数据集上,使用六种不同模态的大量实验表明,X-Fi在人体姿态估计(HPE)与人体活动识别(HAR)任务中均达到领先水平。结果表明,该模型能高效支持多样化的感知应用,推动可扩展多模态感知技术发展。

原文摘要 · Abstract (English)

Human sensing, which employs various sensors and advanced deep learning technologies to accurately capture and interpret human body information, has significantly impacted fields like public security and robotics. However, current human sensing primarily depends on modalities such as cameras and LiDAR, each of which has its own strengths and limitations. Furthermore, existing multi-modal fusion solutions are typically designed for fixed modality combinations, requiring extensive retraining when modalities are added or removed for diverse scenarios. In this paper, we propose a modality-invariant foundation model for all modalities, X-Fi, to address this issue. X-Fi enables the independent or combinatory use of sensor modalities without additional training by utilizing a transformer structure to accommodate variable input sizes and incorporating a novel "X-fusion" mechanism to preserve modality-specific features during multimodal integration. This approach not only enhances adaptability but also facilitates the learning of complementary features across modalities. Extensive experiments conducted on the MM-Fi and XRF55 datasets, employing six distinct modalities, demonstrate that X-Fi achieves state-of-the-art performance in human pose estimation (HPE) and human activity recognition (HAR) tasks. The findings indicate that our proposed model can efficiently support a wide range of human sensing applications, ultimately contributing to the evolution of scalable, multimodal sensing technologies.

多模态感知基础模型姿态估计融合机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。