用扩散模型对齐骨骼与文本特征,提升零样本动作识别准确率
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition

- 将文本特征融入反向扩散过程,引导骨骼特征去噪以实现跨模态对齐
- 在多个数据集上比最新方法提升2.36%至13.05%的准确率
- 适合关注零样本动作识别与跨模态对齐的研究者
在零样本骨骼动作识别(ZSAR)中,对齐骨骼特征与动作标签的文本特征对于准确预测未见动作至关重要。然而,两者之间的模态鸿沟严重限制了模型对未见动作的泛化能力。现有方法多聚焦于直接对齐骨骼与文本潜在空间,但此类空间间的模态差异阻碍了稳健的泛化学习。受扩散模型在多模态对齐(如文生图、文生视频)中的成功启发,本文首次提出一种基于扩散模型的骨骼-文本对齐框架——三元组扩散骨骼-文本匹配(TDSM)。TDSM侧重于扩散模型的跨模态对齐能力而非生成能力:通过将文本特征引入反向扩散过程,使骨骼特征在文本指导下进行去噪,从而构建统一的骨骼-文本潜在空间以实现鲁棒匹配。为增强判别力,引入三元组扩散(TD)损失,促使TDSM纠正错误匹配并拉大不同动作类别的特征距离。实验表明,TDSM显著超越近期最先进方法,在多个基准上取得2.36%至13.05%的性能提升,验证了其在零样本场景下的高精度与强可扩展性。
原文摘要 · Abstract (English)
In zero-shot skeleton-based action recognition (ZSAR), aligning skeleton features with the text features of action labels is essential for accurately predicting unseen actions. ZSAR faces a fundamental challenge in bridging the modality gap between the two-kind features, which severely limits generalization to unseen actions. Previous methods focus on direct alignment between skeleton and text latent spaces, but the modality gaps between these spaces hinder robust generalization learning. Motivated by the success of diffusion models in multi-modal alignment (e.g., text-to-image, text-to-video), we firstly present a diffusion-based skeleton-text alignment framework for ZSAR. Our approach, Triplet Diffusion for Skeleton-Text Matching (TDSM), focuses on cross-alignment power of diffusion models rather than their generative capability. Specifically, TDSM aligns skeleton features with text prompts by incorporating text features into the reverse diffusion process, where skeleton features are denoised under text guidance, forming a unified skeleton-text latent space for robust matching. To enhance discriminative power, we introduce a triplet diffusion (TD) loss that encourages our TDSM to correct skeleton-text matches while pushing them apart for different action classes. Our TDSM significantly outperforms very recent state-of-the-art methods with significantly large margins of 2.36%-point to 13.05%-point, demonstrating superior accuracy and scalability in zero-shot settings through effective skeleton-text matching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。