arXiv:2510.04753cs.CV2025-10被引 2

用人体姿态与动作识别对话中的人,效果优于传统方法。

Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics

  • 分路建模身体姿态和动态变化,用变换器捕捉多尺度运动特征。
  • 姿态信息比动作更关键,融合后识别准确率达98.03%。
  • 适合研究自然交互、跨文化行为分析的学者参考。

本文研究了基于变换器架构在自然面对面对话场景中的人物身份识别性能。我们构建并评估了一个双流框架,分别对133个COCO WholeBody关键点的空间配置和时间运动模式进行建模,数据来自CANDOR对话语料库的一个子集。实验比较了预训练与从头训练的效果,探讨了速度特征的使用,并引入多尺度时间变换器实现分层运动建模。结果表明,领域特定训练显著优于迁移学习,且空间配置比时间动态携带更多判别性信息。空间变换器达到95.74%准确率,多尺度时间变换器达93.90%。特征级融合将性能提升至98.03%,证实姿势与动态信息具有互补性。这些发现展示了变换器在自然交互中人物识别的潜力,并为未来多模态与跨文化研究提供了洞见。

原文摘要 · Abstract (English)

This paper investigates the performance of transformer-based architectures for person identification in natural, face-to-face conversation scenario. We implement and evaluate a two-stream framework that separately models spatial configurations and temporal motion patterns of 133 COCO WholeBody keypoints, extracted from a subset of the CANDOR conversational corpus. Our experiments compare pre-trained and from-scratch training, investigate the use of velocity features, and introduce a multi-scale temporal transformer for hierarchical motion modeling. Results demonstrate that domain-specific training significantly outperforms transfer learning, and that spatial configurations carry more discriminative information than temporal dynamics. The spatial transformer achieves 95.74% accuracy, while the multi-scale temporal transformer achieves 93.90%. Feature-level fusion pushes performance to 98.03%, confirming that postural and dynamic information are complementary. These findings highlight the potential of transformer architectures for person identification in natural interactions and provide insights for future multimodal and cross-cultural studies.

人物识别变换器姿态分析对话理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。