arXiv:2607.17121cs.CV2026-07中稿 · ACII2026 workshop

通过多流骨架建模与弱标签分布学习,提升身体动作中的情绪识别准确率。

Learning Emotion from Motion: Kinetic Multi-Stream Skeleton Modeling with Metadata-Conditioned Weak Label Distributions

论文配图:Learning Emotion from Motion: Kinetic Multi-Stream Skeleton Modeling with Metadata-Conditioned Weak Label Distributions
图 1 · 摘自论文原文
  • 融合旋转、部位感知与元数据条件的弱标签分布学习三分支模型
  • 在DIEM-A任务上准确率提升至36.6%,宏平均F1达35.3%
  • 揭示肢体运动速度和骨骼变化是情绪识别的关键线索

基于骨架的情绪识别仍具挑战性,因情绪表达常依赖细微的动态与关系运动线索,且硬标签难以捕捉相关情绪类别间的模糊性。针对MMAC ACII 2026挑战赛的DIEM-A任务,本文提出一种多分支骨架情绪识别框架,包含基于6D旋转的分支、部位感知的动能多流分支,以及元数据条件化的弱标签分布学习(LDL)分支。各分支独立训练,推理时在概率层面进行集成。在10折留表演者外交叉验证中,该框架将旋转基线的准确率从0.271提升至0.366,宏平均F1从0.252提升至0.353。可解释性分析表明,速度流与骨连接流,以及手臂和腿部区域提供了重要情绪识别线索。

原文摘要 · Abstract (English)

Skeleton-based emotion recognition from body motion remains challenging because emotional expressions are often characterized by subtle dynamic and relational motion cues, and hard labels may not fully capture ambiguity among related emotion categories. For the DIEM-A task in the MMAC ACII 2026 Challenge, we propose a multi-branch skeleton-based emotion recognition framework that combines a 6D rotation-based branch, a part-aware kinetic multi-stream branch, and a metadata-conditioned weak label distribution learning (LDL) branch. The branches are trained independently and fused by a probability-level ensemble at inference time. In 10-fold leave-performer-out cross-validation, the proposed framework improves Accuracy from 0.271 to 0.366 and Macro-F1 from 0.252 to 0.353 over the rotation-based baseline. Explainability ablations show that velocity and bone streams, as well as arm and leg regions, provide important cues for recognizing emotional body motion.

情绪识别骨架建模弱监督动作分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。