无需标签提取步态结构线索,提升识别准确率。
Sketch It Out: Exploring Label-Free Structural Cues for Multimodal Gait Recognition
- 用边缘检测从图像直接提取无标签的结构特征
- 在SUSTech1K上达92.9% Rank-1,在CCPG上达93.1%均值Rank-1
- 适合追求高精度且避免标注成本的步态识别场景
步态识别是安全应用中一种非侵入式生物特征技术,现有方法多依赖轮廓或解析表示。轮廓信息稀疏,丢失内部结构细节,限制区分能力;解析虽补充了部件级结构,但严重依赖上游人体解析器(如标签粒度和边界精度),导致跨数据集性能不稳定,有时甚至劣于轮廓。本文从结构视角重新审视步态表征,定义由边缘密度与监督形式构成的设计空间:轮廓使用稀疏边界边与弱单标签监督,解析则使用更密集的线索与强语义先验。在此空间中,我们识别出一个未充分探索的范式:密集部件级结构但无显式语义标签。为此提出SKETCH作为新视觉模态,通过边缘检测器从RGB图像中无标签地提取高频结构线索(如肢体关节运动与自遮挡轮廓)。进一步证明标签引导解析与无标签草图在语义上解耦、结构上互补。基于此,提出SKETCHGAIT框架,采用分层解耦的多模态结构,包含两个独立流进行模态特异性学习,并设轻量级早期融合分支以捕捉结构互补性。在SUSTech1K与CCPG上的大量实验验证了该模态与框架的有效性:SketchGait在SUSTech1K上实现92.9% Rank-1,在CCPG上实现93.1%平均Rank-1。
原文摘要 · Abstract (English)
Gait recognition is a non-intrusive biometric technique for security applications, yet existing studies are dominated by silhouette- and parsing-based representations. Silhouettes are sparse and miss internal structural details, limiting discriminability. Parsing enriches silhouettes with part-level structures, but relies heavily on upstream human parsers (e.g., label granularity and boundary precision), leading to unstable performance across datasets and sometimes even inferior results to silhouettes. We revisit gait representations from a structural perspective and describe a design space defined by edge density and supervision form: silhouettes use sparse boundary edges with weak single-label supervision, while parsing uses denser cues with strong semantic priors. In this space, we identify an underexplored paradigm: dense part-level structure without explicit semantic labels, and introduce SKETCH as a new visual modality for gait recognition. Sketch extracts high-frequency structural cues (e.g., limb articulations and self-occlusion contours) directly from RGB images via edge-based detectors in a label-free manner. We further show that label-guided parsing and label-free sketch are semantically decoupled and structurally complementary. Based on this, we propose SKETCHGAIT, a hierarchically disentangled multi-modal framework with two independent streams for modality-specific learning and a lightweight early-stage fusion branch to capture structural complementarity. Extensive experiments on SUSTech1K and CCPG validate the proposed modality and framework: SketchGait achieves 92.9% Rank-1 on SUSTech1K and 93.1% mean Rank-1 on CCPG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。