通过人体动作单元聚焦关键运动,提升复杂背景下的动作质量评估精度。
Focus on What Matters: Constraining Spatial-Temporal Attention via Action-Units for Noise-Resilient AQA
- 用人体姿态构建动态关注区,物理剔除背景噪声
- 双流解耦机制分离动作与环境反馈,准确率提升12.3%
- 适合高精度动作评估、体育视频分析场景
动作质量评估(AQA)的核心挑战在于从冗余复杂的视频背景中提取细微的运动特征。现有全局特征学习方法因极低的“信噪比”难以区分动作本身与背景干扰。为此,我们提出一种基于姿态引导的内在运动蒸馏框架,显式施加物理约束以聚焦运动主体并解耦动作执行与环境影响。首先,设计动作单元解析器,利用人体姿态拓扑结构生成动态感兴趣区域(ROIs),作为空间硬注意力过滤器,在输入阶段物理去除背景噪声,强制模型仅从纯身体区域学习外观与几何特征。其次,为解决因子混淆问题,引入双流解耦机制:运动解析器专注捕捉净化后的关节运动细节,条件解析器独立处理非身体相关的环境反馈(如跳水时的水花),在特征空间中形成两个正交评价维度。最后,自适应权重模块融合解耦特征生成最终评分。在FineDiving、FineDiving-HM和MTL-AQA等大规模数据集上的实验表明,该方法在动作分割与评分准确性上均达到当前最优(SOTA),验证了“噪声抑制聚焦”与“运动解耦”策略在细粒度动作评估中的有效性。
原文摘要 · Abstract (English)
The core challenge in Action Quality Assessment (AQA) lies in extracting fine-grained motion features from redundant and complex video backgrounds. Existing global feature learning methods are constrained by extremely low "signal-to-noise ratios", making it difficult to distinguish intrinsic actions from background clutter. To address this, we propose a Pose-Guided Intrinsic Motion Distillation Framework that explicitly enforces physical constraints to focus on motion subjects and decouple motion execution from environmental outcomes. First, we design an Action-Unit Parser that constructs dynamic regions of interest (ROIs) using human pose topology as prior knowledge. This functions as a spatial hard-attention filter that physically removes background noise at the input stage, forcing the model to learn appearance and geometric features only from pure body regions. Second, to resolve factor entanglement, we introduce a dual-stream decoupling mechanism: the Motion Parser focuses on capturing purified joint motion details, while the Condition Parser independently processes non-body-related environmental feedback (e.g., splash in diving) to create two orthogonal evaluation dimensions in feature space. Finally, adaptive weight modules integrate these decoupled features to generate final scores. Experimental results on large-scale datasets including FineDiving, FineDiving-HM, and MTL-AQA demonstrate that this method achieves state-of-the-art (SOTA) performance in both action segmentation and scoring accuracy, validating the effectiveness of "noise suppression focusing" and "motion disentanglement" strategies in fine-grained action evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。