arXiv:2608.13092cs.CV2026-08中稿 · ACM MM2026

提出统一框架,融合RGB与事件视频信息,提升跨摄像头行人检索性能。

Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification

论文配图:Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification
图 1 · 摘自论文原文
  • 设计记忆增强主干网络,保持模态特异性身份原型。
  • 引入提示感知时空变换器,联合建模空间与时间特征。
  • 分层多模态融合策略,兼顾全局与局部判别性信息,适合多模态行人识别任务。

RGB-事件视频行人重识别(RE-VReID)旨在通过互补的RGB视频和事件流,在非重叠摄像头间检索特定行人。现有方法常将空间与时间建模解耦,限制了交互;且全局级多模态融合难以充分挖掘细粒度判别线索。为此,我们提出Paths,一个统一的框架,包含时空建模与分层多模态融合机制。首先设计记忆增强主干网络(MAB),维护模态特异性身份原型,实现稳定的模态内表征学习。其次提出提示感知时空变换器(PST),在统一的Transformer中联合建模空间与时间线索。最后引入分层多模态融合(HMF),在全局与局部层面融合RGB与事件特征。在EvReID、MARS和iLIDS-VID三个公开数据集上的大量实验验证了方法的有效性。代码已开源:https://github.com/Reflection0427/Paths。

原文摘要 · Abstract (English)

RGB-Event Video Person Re-Identification (RE-VReID) aims to retrieve specific person across non-overlapping cameras with complementary RGB videos and event streams. However, existing methods often decouple spatial and temporal modeling, which limits their interaction. In addition, global-level RGB-Event fusion fails to fully exploit fine-grained discriminative cues. To address these issues, we propose Paths, a unified framework with spatio-temporal modeling and hierarchical multi-modal fusion for RE-VReID. Specifically, we first design a Memory-Augmented Backbone (MAB) to maintain modality-specific identity prototypes for stable intra-modal representation learning. Then, we propose a Prompt-aware Spatio-temporal Transformer (PST) to jointly model spatial and temporal cues within a unified Transformer. Finally, we introduce a Hierarchical Multi-modal Fusion (HMF) to integrate RGB and event features at global and local levels. With these modules, our framework can learn robust and discriminative representations for RE-VReID. Extensive experiments on three public RE-VReID benchmarks including EvReID, MARS and iLIDS-VID, demonstrate the effectiveness of our proposed method. The code is available at https://github.com/Reflection0427/Paths.

行人重识别多模态融合事件相机时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。