arXiv:2410.21108cs.CV2024-10中稿 · WACV 2025被引 5

用激光雷达引导视觉文本,提升群体活动识别准确率

LiGAR: LiDAR-Guided Hierarchical Transformer for Multi-Modal Group Activity Recognition

  • 以激光雷达结构为骨架,分层融合视觉与文本信息
  • 在JRDB-PAR上F1提升10.6%,NBA数据集准确率提高5.9%
  • 无需激光雷达也能保持性能,适合复杂场景应用

群体活动识别因多主体交互复杂而具挑战性。本文提出LiGAR,一种基于激光雷达引导的分层变压器模型,用于多模态群体活动识别。该模型利用激光雷达数据作为结构基础,指导视觉与文本信息处理,有效应对遮挡与复杂空间布局。框架包含多尺度激光雷达变换器、跨模态引导注意力和自适应融合模块,实现不同语义层级的多模态数据融合。其分层结构可捕捉从个体动作到场景动态的多层次群体行为。在JRDB-PAR、Volleyball和NBA数据集上的实验表明,LiGAR性能优越,相较现有方法在JRDB-PAR上F1-score提升最高达10.6%,在NBA数据集上平均分类准确率提升5.9%。值得注意的是,即使推理时无激光雷达数据,模型仍保持高精度,展现出强适应性。消融实验验证了各组件的关键贡献及多模态、多尺度方法的有效性。

原文摘要 · Abstract (English)

Group Activity Recognition (GAR) remains challenging in computer vision due to the complex nature of multi-agent interactions. This paper introduces LiGAR, a LIDAR-Guided Hierarchical Transformer for Multi-Modal Group Activity Recognition. LiGAR leverages LiDAR data as a structural backbone to guide the processing of visual and textual information, enabling robust handling of occlusions and complex spatial arrangements. Our framework incorporates a Multi-Scale LIDAR Transformer, Cross-Modal Guided Attention, and an Adaptive Fusion Module to integrate multi-modal data at different semantic levels effectively. LiGAR's hierarchical architecture captures group activities at various granularities, from individual actions to scene-level dynamics. Extensive experiments on the JRDB-PAR, Volleyball, and NBA datasets demonstrate LiGAR's superior performance, achieving state-of-the-art results with improvements of up to 10.6% in F1-score on JRDB-PAR and 5.9% in Mean Per Class Accuracy on the NBA dataset. Notably, LiGAR maintains high performance even when LiDAR data is unavailable during inference, showcasing its adaptability. Our ablation studies highlight the significant contributions of each component and the effectiveness of our multi-modal, multi-scale approach in advancing the field of group activity recognition.

群体识别激光雷达多模态分层模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。