区分日常动作靠姿势,还原动作过程需动态信息。
Classifying daily activities needs posture, reconstructing them needs motion

- 用三种方法分解动作:临时运动基元、勒让德系数、自编码器嵌入。
- 姿势特征准确率最高,9个关键关节决定动作分类。
- 动态信息对动作重建至关重要,仅靠姿势会生成僵化结果。
人类能轻松从复杂视觉输入中识别动作,但其依赖的刺激信息尚不明确。本研究基于MoVi数据集中的16种日常活动视频,系统比较了三种动作分析策略:时间运动基元(TMPs)、勒让德多项式系数和自编码器隐向量。结果显示,勒让德系数与TMPs在动作分类中表现最佳,其次为自编码器。识别出两类判别性特征:一是身体整体姿势(平均空间构型),可有效区分不同动作;二是9个最具预测性的关键关节。有趣的是,高分类精度并不等于良好动作重建:TMPs能保留时间动态,生成自然运动;而勒让德系数仅保留平均姿势,导致动作看起来僵化。这揭示了动作信息组织的分离性——静态姿势足以判断动作类型,但动态信息才是重建动作过程的关键。该发现有助于理解视觉系统快速识别动作的机制,并提示姿势特征可用于临床动作筛查,而动态信息对生成任务仍不可或缺。
原文摘要 · Abstract (English)
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。