通过肢体动作推断心理状态,无需问卷且保护隐私
From Actions to Kinesics: Extracting Human Psychological States through Bodily Movements
- 用3D骨骼数据结合图卷积与CNN模型识别肢体语言含义
- 在DUET数据集上实现高精度心理状态推断,效果优于传统方法
- 适合研究人机交互、心理学或强化学习中的行为建模者
理解人类与建筑环境之间的动态关系是环境心理学到强化学习等多个领域的重要挑战。建模这类互动的核心障碍在于无法以通用且保护隐私的方式捕捉人类心理状态。传统方法依赖理论模型或问卷调查,存在范围有限、静态且耗时的问题。本文提出一种基于运动学的识别框架,直接从3D骨骼关节数据中推断人体活动的交际功能(即运动学)。该框架结合时空图卷积网络(ST-GCN)与卷积神经网络(CNN),利用迁移学习避免手动定义物理动作与心理类别的映射关系。方法在保留用户匿名性的同时,揭示了反映认知与情绪状态的潜藏运动模式。在双人用户参与度数据集(DUET)上的实验表明,该方法可实现可扩展、准确且以人为本的行为建模,为提升强化学习驱动的人-环境交互模拟提供了新路径。
原文摘要 · Abstract (English)
Understanding the dynamic relationship between humans and the built environment is a key challenge in disciplines ranging from environmental psychology to reinforcement learning (RL). A central obstacle in modeling these interactions is the inability to capture human psychological states in a way that is both generalizable and privacy preserving. Traditional methods rely on theoretical models or questionnaires, which are limited in scope, static, and labor intensive. We present a kinesics recognition framework that infers the communicative functions of human activity -- known as kinesics -- directly from 3D skeleton joint data. Combining a spatial-temporal graph convolutional network (ST-GCN) with a convolutional neural network (CNN), the framework leverages transfer learning to bypass the need for manually defined mappings between physical actions and psychological categories. The approach preserves user anonymity while uncovering latent structures in bodily movements that reflect cognitive and emotional states. Our results on the Dyadic User EngagemenT (DUET) dataset demonstrate that this method enables scalable, accurate, and human-centered modeling of behavior, offering a new pathway for enhancing RL-driven simulations of human-environment interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。