arXiv:2609.02510cs.CVcs.HC2026-09

通过正交模型集成提升无演员依赖的情绪识别,效果显著且可解释。

Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition

论文配图:Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition
图 1 · 摘自论文原文
  • 用11个误差模式正交的模型集成,提升分类性能。
  • 在留一人外验证下,宏平均F1达36.80%,比基线高11.07个百分点。
  • 首次用可解释性分析证明情绪判断依赖身体区域运动特征,而非传统运动学。

我们研究在留演员外(LPO)评估设置下的仅基于骨架动作的12类表演情绪分类,这是一个困难且欠定的问题:随机猜测准确率为8.3%,协议匹配的重现STGCN++基线仅达25.73 ± 4.03%的宏平均F1。我们发现,可靠提升并非来自新架构,而是通过组合11个误差模式正交的模型实现:在标注训练演员上的10折LPO交叉验证中,等权重对数均值集成达到36.80 ± 4.00%每折宏平均F1,比相同数据划分的基线高出11.07个百分点(相对提升43%)。核心贡献是一套经验证的可解释性分析:对强模型成员,局部遮蔽与反事实编辑表明其决策依赖于运动相关的身体区域证据,且该区域显著性与基于规则的拉班动作分析(LMA)属性相关性(皮尔逊等级相关系数rho = +0.500)远高于经典运动学(rho = +0.033),约为15倍;该关联在提交的11模型集成本身也成立(rho = +0.517);审计为事后分析,无需重新训练。同一套方法亦准确揭示:时间维度上的显著性分布广泛而非集中。

原文摘要 · Abstract (English)

We study body-only, 12-class acted-emotion classification from skeleton motion under leave-performer-out (LPO) evaluation, a hard, underdetermined setting: chance is 8.3%, and a protocol-matched reproduced STGCN++ baseline reaches only 25.73 +/- 4.03% Macro-F1. We show that reliable gains come not from a new architecture but from combining eleven models with orthogonal error modes: under 10-fold LPO cross-validation on the labeled training performers, an equal-weight logit-mean ensemble reaches 36.80 +/- 4.00% per-fold Macro-F1, a protocol-matched +11.07 pp (+43% relative) over the same-split reproduced baseline. Our central contribution is a tested explanation suite: for a strong ensemble member, part-masking and counterfactual edits show (rather than assert) that its decisions depend on motion-grounded body-region evidence, and this region saliency aligns with rule-based Laban Movement Analysis (LMA) attributes far more than with classical kinematics: region-level saliency-LMA Spearman rho = +0.500 versus +0.033, roughly 15x, and the alignment holds for the submitted 11-way ensemble itself at rho = +0.517; the audit is post hoc and needs no retraining. The same suite faithfully reports a negative: within-window temporal saliency is diffuse rather than localized.

情绪识别可解释性模型集成动作分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。