用文字指导分离身份与动作特征,提升跨活动身份识别准确率。
DisenQ: Disentangling Q-Former for Activity-Biometrics
- 通过语言引导的查询变换器,分离身份、动作和非身份特征。
- 在三个活动视频基准上达到当前最优性能,跨场景泛化能力强。
- 无需依赖姿态或轮廓等易错视觉数据,适合真实复杂场景应用。
本文研究跨活动身份生物识别问题,即在多种动作下识别个体。传统方法因身份线索与运动动态、外观变化纠缠而困难重重。现有依赖姿态或轮廓等额外视觉信息的方法常受提取误差影响。为此,我们提出一种多模态语言引导框架,以结构化文本监督替代对额外视觉数据的依赖。核心是提出 extbf{DisenQ}( extbf{Disen}tangling extbf{Q}-Former),一个统一的查询变换器,通过语言引导实现身份、运动和非身份特征的解耦,确保身份特征独立于外观与动作变化,避免误识。我们在三个基于活动的视频基准上评估,取得当前最优性能。此外,在传统视频识别基准上也展现出良好泛化能力,验证了框架的有效性。
原文摘要 · Abstract (English)
In this work, we address activity-biometrics, which involves identifying individuals across diverse set of activities. Unlike traditional person identification, this setting introduces additional challenges as identity cues become entangled with motion dynamics and appearance variations, making biometrics feature learning more complex. While additional visual data like pose and/or silhouette help, they often struggle from extraction inaccuracies. To overcome this, we propose a multimodal language-guided framework that replaces reliance on additional visual data with structured textual supervision. At its core, we introduce \textbf{DisenQ} (\textbf{Disen}tangling \textbf{Q}-Former), a unified querying transformer that disentangles biometrics, motion, and non-biometrics features by leveraging structured language guidance. This ensures identity cues remain independent of appearance and motion variations, preventing misidentifications. We evaluate our approach on three activity-based video benchmarks, achieving state-of-the-art performance. Additionally, we demonstrate strong generalization to complex real-world scenario with competitive performance on a traditional video-based identification benchmark, showing the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。