解决微动作识别中长尾分布与模糊样本难题,提升细粒度分类准确率。
SoftRerank: Hierarchical Soft Fusion with Candidate-Label Reranking for Long-Tailed Micro-Action Recognition

- 分层软融合+候选标签重排序,增强细粒度判断能力。
- 在MA-52数据集上达到79.99%的F1均值,夺冠ACM Multimedia 2026挑战赛。
- 适合处理弱视觉变化、类别混淆严重的微动作识别任务。
微动作是细微、低强度的非语言行为,能反映人的精细状态(如情绪与意图)。由于其持续时间短、视觉变化微弱且不同类别间运动模式相似,识别难度大。本文提出一种细粒度微动作识别方法,结合InternVideo2.5全量微调、分层软融合与轻量级候选标签重排序器。针对MA-52数据集的长尾标签分布,采用类别平衡采样和逆频率加权以降低高频类影响。端到端微调InternVideo2.5,并在共享视频表征上添加粗粒度与组条件细粒度分类头,提升粗粒度与细粒度预测的一致性。对模糊样本,重排序器利用困难样本与视频-标签匹配,聚焦易混淆的细粒度动作。实验验证该方法在MA-52上取得79.99%的F1-mean,位列第3届ACM Multimedia 2026微动作分析大赛第一。
原文摘要 · Abstract (English)
Micro-actions are subtle, low-intensity non-verbal behaviors that provide cues to fine-grained human states, including emotions and intentions. Recognizing them remains difficult because they are brief, contain weak visual changes, and often exhibit similar motion patterns across categories. This paper addresses these challenges with a fine-grained micro-action recognition method that combines full fine-tuning of InternVideo2.5, hierarchical soft fusion, and a lightweight candidate-label reranker. For the long-tailed label distribution in MA-52, we use class-balanced sampling and inverse-frequency reweighting to reduce the effect of frequent classes during training. We fine-tune InternVideo2.5 end to end and attach coarse and group-conditional fine-grained classification heads to the shared video representation, improving the consistency between coarse and fine predictions. For ambiguous samples, the candidate-label reranker uses hard samples and video-label matching to focus on easily confused fine-grained actions. Experiments validate the proposed method, which achieves a 79.99% F1-mean on MA-52 and ranks first in the 3rd Micro-Action Analysis Grand Challenge at ACM Multimedia 2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。