用文字描述辅助分离步态特征,提升跨模态识别准确率
Text-guided Feature Disentanglement for Cross-modal Gait Recognition

- 用大模型生成步态文字描述,作为语义锚点指导特征分离
- 在SUSTech1K和FreeGait上达到新最好效果,准确率超现有方法
- 适合做多模态生物识别、尤其是激光雷达与摄像头融合场景
步态识别通过行走姿态识别个体,在远距离、非侵入场景中具有优势。但真实场景常涉及异构传感模态(如激光雷达与可见光相机),导致2D视频与3D点云序列间存在显著模态差距,使激光雷达-相机跨模态步态识别(LCCGR)成为关键挑战。为此,我们提出TCFDNet:一种文本引导的跨模态特征解耦网络,利用模态感知的文本先验作为语义锚点,引导学习解耦的共享特征。具体地,我们使用大语言模型构建步态模态文本词典(GMTD),生成多模态与视角下的丰富语义描述;基于CLIP的多粒度特征编码器将视觉与文本特征对齐至统一视觉-语言空间;此外,文本引导特征解耦(TFD)模块选取topk匹配文本描述,通过残差分解与正交性约束重构模态特定特征并提取共享特征。为增强共享特征鲁棒性,设计特征稳定性增强(FSE)模块,建模空间与通道相关性;同时引入跨模态补丁交换策略提升泛化能力。在SUSTech1K与FreeGait数据集上的大量实验表明,TCFDNet实现新的最优性能,验证了各模块有效性。
原文摘要 · Abstract (English)
Gait recognition is a biometric technique that identifies individuals based on their walking patterns, offering advantages in long-range, non-intrusive scenarios. However, real-world scenarios often involve heterogeneous sensing modalities such as LiDAR and RGB cameras, making LiDAR-Camera Cross-modal Gait recognition (LCCGR) a critical yet challenging task due to the substantial modality gap between 2D videos and 3D point cloud sequences. To address this challenge, we propose TCFDNet, a Text-guided Cross-modal Feature Disentanglement Network, which leverages modality-aware textual priors as semantic anchors to guide the learning of disentangled modality-shared representations. Specifically, we construct a Gait Modality Text Dictionary (GMTD) using large language models to generate rich semantic descriptions of gait across modalities and viewpoints. A CLIP-based Multi-grained Feature Encoder then aligns visual and textual features within a unified vision-language space. Furthermore, the Text-guided Feature Disentanglement (TFD) module selects the topk matched textual descriptions to reconstruct modality-specific representations and derive modality-shared features via residual decomposition and orthogonality constraints. To mitigate the fragility of the disentangled shared features, we propose a Feature Stability Enhancement (FSE) module, which models spatial and channel-wise correlations to improve feature robustness. In addition, a cross-modal patch exchange strategy is introduced to further improve generalization. Extensive experiments on SUSTech1K and FreeGait datasets demonstrate that TCFDNet achieves new state-of-the-art results and validate the effectiveness of the proposed modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。