用自然语言指令精准检索特定人不同步态,解决身份漂移问题。
Who Remains, What Changes: Identity Anchored Composed Gait Retrieval

- 基于参考步态和语言指令生成新步态,通过身份锚定防止认错人。
- 在两个新构建数据集上达到72.38%和83.61%的R@1准确率。
- 适合需要交互式步态检索的安防与个性化识别场景。
步态识别虽取得显著进展,但现有方法仍局限于刚性视觉匹配,忽视自然语言指令在交互检索中的潜力。本文提出组合步态检索(CoGR)新任务,即根据参考步态序列和自然语言修改指令检索目标步态。为弥补数据缺失,我们利用大视觉语言模型设计自动化标注流程,构建首个步态-语言数据集:Language-Augmented CCPG 和 Language-Augmented CASIA-B。在此基础上,提出身份锚定的组合框架ComposeGait,其部分感知身份适配器(PIA)将多帧、部位感知的身份证据聚合为样本专属身份令牌,并注入共享Q-Former双分支以保持身份一致,同时排除身份令牌输出至最终检索嵌入。联合身份与任务自适应的组合检索目标实现端到端优化。在两个基准上评估显示,ComposeGait 在 Language-Augmented CCPG 上达到72.38% R@1,Language-Augmented CASIA-B 上达83.61%,建立强基线。数据与代码将公开。
原文摘要 · Abstract (English)
Gait recognition has achieved remarkable progress, yet existing methods remain confined to rigid visual matching and often overlook the potential of natural language instructions for interactive retrieval. In this paper, we introduce Composed Gait Retrieval (CoGR), a novel task that retrieves a target gait sequence based on a reference sequence and a natural language modification query. To address the absence of existing datasets for this task, we design an automated annotation pipeline powered by large vision-language models (VLMs) to construct the first gait-language datasets: Language-Augmented CCPG and Language-Augmented CASIA-B. Building on this, we propose ComposeGait, an identity-anchored composition framework designed to prevent the identity drift that arises when generic composed retrieval follows the instruction but returns the wrong person. Its Part-aware Identity Adapter (PIA) aggregates multi-frame, part-aware identity evidence into a sample-specific ID token. We inject the ID tokens into both branches of a shared Q-Former to preserve identity, while excluding the ID-token outputs from the final retrieval embeddings. Joint identity and task-adapted composed-retrieval objectives optimize this space end to end. We evaluate ComposeGait on both benchmarks and show that it achieves the best R@1 among the compared methods, reaching 72.38% on Language-Augmented CCPG and 83.61% on Language-Augmented CASIA-B. These results establish ComposeGait as a strong baseline for CoGR. The datasets and code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。