arXiv:2606.00153cs.CVcs.AI2026-06中稿 · ICML

通过潜在扩散空间实现2D/3D步态轨迹对齐,提升跨模态识别精度。

DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion

论文配图:DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
图 1 · 摘自论文原文
  • 在潜在扩散空间中进行轨迹级对齐,而非仅对齐最终特征
  • 在SUSTech1K和FreeGait上达到当前最优性能
  • 生成对齐与分类解耦,推理效率高,适合实际部署

跨模态2D-3D步态识别受限于2D轮廓与3D LiDAR范围图表示之间的固有域差异。现有方法仅对齐最终嵌入,我们提出DiffCrossGait,将跨模态匹配重新定义为身份相关潜在扩散空间中的轨迹级对齐,而非假设2D与3D观测完全等价。通过在潜在空间中使用共享高斯噪声驱动双模态,实现在生成演化过程中的连续对齐。引入三阶段对齐策略,利用不同噪声强度分别强化身份锚定、动态一致性与跨模态结构可恢复性,从而约束双模态共享去噪动态与瓶颈结构,促进模态不变的步态特征。关键在于,该框架将生成对齐与判别主干解耦,扩散机制仅作为训练目标,避免迭代去噪的计算开销,保障高推理效率。在SUSTech1K与FreeGait基准上的大量实验表明,DiffCrossGait取得最先进性能。

原文摘要 · Abstract (English)

Cross-modal 2D-3D gait recognition is impeded by inherent domain discrepancies between 2D silhouette and 3D LiDAR range-view representations. While prior methods align only final embeddings, we propose DiffCrossGait, which reformulates cross-modal matching as trajectory-level alignment in an identity-relevant latent diffusion space, rather than assuming full equivalence between 2D and 3D observations. By driving both modalities with shared Gaussian noise within a latent space, we enable continuous alignment throughout the generative evolution. We introduce a Tri-Phase Alignment Strategy that exploits varying noise intensities to enforce identity anchoring, dynamics consistency, and cross-modal structural recoverability, thereby constraining both modalities to share denoising dynamics and bottleneck structure, which promotes modality-invariant gait features. Crucially, our framework decouples generative alignment from the discriminative backbone; the diffusion mechanism serves exclusively as a training objective, ensuring high inference efficiency by eliminating the computational overhead of iterative denoising. Extensive experiments on the SUSTech1K and FreeGait benchmarks demonstrate that DiffCrossGait achieves state-of-the-art performance.

步态识别跨模态扩散模型特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。