研究发现视频模型常依赖外观而非运动特征,抑制外观后模型才学会用独特动作识别人
Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study

- 构建120名篮球运动员罚球动作数据集,控制动作和拍摄差异
- 模型在外观可见时主要靠脸和球衣识人,外观被遮蔽后转而关注细微动作
- 屏蔽外观后模型更鲁棒,且能捕捉个体特有动作模式,适合身份识别场景
身份识别传统上依赖静态外观特征。然而,一致且个体特有的运动动态可提供互补且更鲁棒的标识,尤其在外观弱或变化时。本研究通过系统性诊断实验,提出BALLER120——一个由120名专业篮球运动员完成罚球动作的受控基准数据集。该数据集聚焦相同多阶段动作,减少动作层面差异和与身份相关的采集偏差,实现对个体运动学模式的精细分析。结果显示,现代视频模型虽能从RGB视频中准确预测身份,但常依赖人脸、球衣等静态外观线索,即使存在丰富的运动信息。令人惊讶的是,当外观被抑制为仅轮廓或骨架输入时,同一模型架构转向学习运动微模式(如脚部落点、肘部弯曲)。尽管视觉信息减少,这些表示仍达到相当准确率,并对外观变化更具鲁棒性。定性分析进一步表明,外观抑制模型会关注个体间独特的运动模式。总体而言,本研究证明个体特异性运动签名存在、具有信息量且可学习,但现代视频模型可能因偏好更简单的静态捷径而忽略它们,除非显式抑制外观线索。
原文摘要 · Abstract (English)
Identity recognition (e.g., person, animal re-identification) has traditionally relied heavily on static appearance cues. Yet motion--consistent, individual-specific dynamics--can provide a complementary and potentially more robust signature, especially when appearance is weak or variable. This raises a fundamental question: when identity-specific motion cues are clearly present, to what extent do modern video models use them for recognition? To investigate this question, we conduct a systematic diagnostic study and introduce BALLER120, a controlled benchmark of 120 professional basketball players performing free-throws. By focusing on the same multi-phase action across individuals, BALLER120 reduces action-level variation and identity-correlated acquisition biases, enabling fine-grained analysis of identity-specific kinematic patterns. We find that modern video models can predict identity accurately from RGB videos, but often rely on static appearance cues such as faces and jersey regions, even when informative motion cues are available. Strikingly, when appearance is suppressed through silhouette-only or skeleton-only inputs, the same model architectures shift toward motion micro-patterns (e.g., foot placement and elbow bending). Despite containing less visual information, appearance-suppressed representations achieve competitive accuracy and stronger robustness to appearance shifts. Our qualitative analyses further show that appearance-suppressed models attend to distinctive motion patterns across individuals. Overall, our study demonstrates that identity-specific motion signatures are present, informative, and learnable, but modern video models may overlook them in favor of easier static shortcuts unless appearance cues are explicitly suppressed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。