arXiv:2508.11218cs.CVcs.LG2025-08被引 4

解决自动驾驶中行人重识别的多模态缺失问题

A CLIP-based Uncertainty Modal Modeling (UMM) Framework for Pedestrian Re-Identification in Autonomous Driving

  • 基于CLIP构建轻量级不确定性模态建模框架
  • 在模态缺失下仍保持高识别准确率和低计算开销
  • 适合资源受限的车载智能感知系统

行人重识别(ReID)是智能感知系统的关键技术,尤其在自动驾驶中,车载摄像头需实时跨视角、跨时间识别行人以支持安全导航与轨迹预测。然而,RGB、红外、草图或文本描述等输入模态存在不确定性或缺失,给传统ReID方法带来挑战。尽管大规模预训练模型具备强大的多模态语义建模能力,但其计算开销限制了在资源受限环境中的部署。为此,我们提出一种轻量级不确定性模态建模(UMM)框架,融合多模态标记映射器、合成模态增强策略与跨模态线索交互学习器,实现统一特征表示,缓解模态缺失影响,并提取不同数据类型间的互补信息。此外,UMM利用CLIP的视觉-语言对齐能力,高效融合多模态输入,无需大量微调。实验表明,该框架在不确定模态条件下展现出强鲁棒性、良好泛化能力与高效计算性能,为自动驾驶场景下的行人重识别提供了可扩展且实用的解决方案。

原文摘要 · Abstract (English)

Re-Identification (ReID) is a critical technology in intelligent perception systems, especially within autonomous driving, where onboard cameras must identify pedestrians across views and time in real-time to support safe navigation and trajectory prediction. However, the presence of uncertain or missing input modalities--such as RGB, infrared, sketches, or textual descriptions--poses significant challenges to conventional ReID approaches. While large-scale pre-trained models offer strong multimodal semantic modeling capabilities, their computational overhead limits practical deployment in resource-constrained environments. To address these challenges, we propose a lightweight Uncertainty Modal Modeling (UMM) framework, which integrates a multimodal token mapper, synthetic modality augmentation strategy, and cross-modal cue interactive learner. Together, these components enable unified feature representation, mitigate the impact of missing modalities, and extract complementary information across different data types. Additionally, UMM leverages CLIP's vision-language alignment ability to fuse multimodal inputs efficiently without extensive finetuning. Experimental results demonstrate that UMM achieves strong robustness, generalization, and computational efficiency under uncertain modality conditions, offering a scalable and practical solution for pedestrian re-identification in autonomous driving scenarios.

行人重识别多模态学习自动驾驶CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。