用注意力对齐视觉与音乐,提升长视频动作质量评估准确率
Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment
- 通过多模态注意力机制显式对齐视觉与音频特征
- 在RG和Fis-V数据集上超越现有方法,显著提升评估精度
- 适合艺术体操、花样滑冰等需同步音乐的动作评估场景
长期动作质量评估(AQA)旨在评估持续数分钟视频中人类活动的质量,对艺术体操、花样滑冰等项目自动评分至关重要。现有方法分为两类:仅依赖视觉特征的单模态方法,难以建模音乐等多模态线索;以及采用简单特征级对比融合的多模态方法,忽视深层跨模态协作与时间动态。为此,本文提出长时多模态注意力一致性网络(LMAC-Net),引入多模态注意力一致性机制,实现视觉与音频信息的稳定融合与表征增强。具体包括:多模态局部查询编码模块捕捉时序语义与跨模态关系,两级评分机制提供可解释结果,并结合注意力与回归损失联合优化多模态对齐与分数融合。在RG与Fis-V数据集上的实验表明,LMAC-Net显著优于现有方法,验证了其有效性。
原文摘要 · Abstract (English)
Long-term action quality assessment (AQA) focuses on evaluating the quality of human activities in videos lasting up to several minutes. This task plays an important role in the automated evaluation of artistic sports such as rhythmic gymnastics and figure skating, where both accurate motion execution and temporal synchronization with background music are essential for performance assessment. However, existing methods predominantly fall into two categories: unimodal approaches that rely solely on visual features, which are inadequate for modeling multimodal cues like music; and multimodal approaches that typically employ simple feature-level contrastive fusion, overlooking deep cross-modal collaboration and temporal dynamics. As a result, they struggle to capture complex interactions between modalities and fail to accurately track critical performance changes throughout extended sequences. To address these challenges, we propose the Long-term Multimodal Attention Consistency Network (LMAC-Net). LMAC-Net introduces a multimodal attention consistency mechanism to explicitly align multimodal features, enabling stable integration of visual and audio information and enhancing feature representations. Specifically, we introduce a multimodal local query encoder module to capture temporal semantics and cross-modal relations, and use a two-level score evaluation for interpretable results. In addition, attention-based and regression-based losses are applied to jointly optimize multimodal alignment and score fusion. Experiments conducted on the RG and Fis-V datasets demonstrate that LMAC-Net significantly outperforms existing methods, validating the effectiveness of our proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。