arXiv:2605.08945cs.CV2026-05

通过多级线索精炼,提升长时多模态动作质量评估的准确性

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

论文配图:MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment
图 1 · 摘自论文原文
  • 分三层次重构动作质量线索:模态内解耦、跨模态动态互补检索、阶段式融合
  • 在体操和滑雪数据集上,皮尔逊相关系数与预测误差均达最优或次优
  • 适合需要精准捕捉长期动作质量变化的研究者与应用开发者

长时多模态动作质量评估(AQA)通过分析数分钟的音视频序列,挖掘有判别性的质量线索以预测评分。现有方法通常使用单一时间编码器建模整个序列,并通过直接对齐或拼接融合模态特征,导致关键线索被全局趋势掩盖、模态冗余削弱,且在一次性评分映射中失真。为此,本文将长时多模态AQA重新定义为质量线索组织问题,提出MLCR多级线索精炼框架。该框架在三个层面组织质量证据:模态内表示、跨模态交互与阶段式聚合。具体地,模态内解耦编码器(IMDE)在保留模态特性的前提下,同时优化全局时间上下文与局部频率细节;跨模态动态互补感知检索模块(CMDCR)基于动态融合状态,逐步检索增量证据并抑制冗余响应;阶段式多模态融合(SMI)块逐步累积模态内与跨模态线索,持续优化融合表示。在体操(Rhythmic Gymnastics)和滑冰(Fis-V)数据集上的实验表明,MLCR在斯皮尔曼相关系数与预测误差上均达到最佳或第二佳表现,验证了其有效性与鲁棒性。

原文摘要 · Abstract (English)

Long-term multimodal action quality assessment (AQA) evaluates action execution in several-minute audiovisual sequences by mining discriminative quality cues for score prediction. Existing multimodal methods usually model entire sequences with a single temporal encoder and fuse modality features by direct alignment or concatenation, causing key cues to be obscured by global trends, weakened by modal redundancy, and distorted during one-shot score mapping. To address this issue, we reformulate long-term multimodal AQA as a quality cue organization problem and propose MLCR, a multi-level cue refinement framework. MLCR organizes quality evidence at three levels: intra-modal representation, cross-modal interaction, and stage-wise aggregation. Specifically, the intra-modal decoupling encoder (IMDE) preserves modality identity while refining global temporal context and local frequency details. The cross-modal dynamic complementarity-aware retrieval (CMDCR) module retrieves incremental evidence conditioned on the evolving fused state and suppresses redundant responses. The stage-wise multimodal integration (SMI) block progressively accumulates intra-modal and cross-modal cues to refine the fused representation. Experiments on the Rhythmic Gymnastics and Fis-V datasets show that MLCR achieves the best or second-best performance in both Spearman correlation and prediction error, demonstrating its effectiveness and robustness.

动作评估多模态长时序列线索精炼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。