arXiv:2606.07355cs.CV2026-06被引 4

分离时空特征建模,提升微手势在线识别精度

Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition

论文配图:Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition
图 1 · 摘自论文原文
  • 拆分时空分支,用轻量卷积分别捕捉动态与位置特征
  • 在挑战赛中取得0.43808的F1分数,排名第一
  • 自适应增强策略缓解类别分布不均问题,适合实际应用

微手势在线识别旨在对未剪辑视频中的细微动作进行时间定位和分类。由于其持续时间极短、运动幅度小且视觉线索模糊,提取具有判别力的时空表征仍具挑战。现有参数高效适配器通常采用单一分支联合建模时空信息,难以捕捉微手势的精细模式。为此,本文提出空间-时间解耦适配器,通过轻量级深度卷积将视频适配分解为独立的时间与空间分支。此外,为缓解基准数据集固有的长尾类别分布问题,引入自适应软平衡增强方法,根据类别稀有度与学习难度动态分配增强强度,无需人工设定阈值。所提方法在第四届EI-MiGA-IJCAI挑战赛第2赛道中取得0.43808的F1分数,排名第一。

原文摘要 · Abstract (English)

Micro-gesture online recognition aims to temporally localize and classify subtle gestures in untrimmed videos. Owing to their extremely short duration, low motion amplitude, and ambiguous visual cues, capturing discriminative spatiotemporal representations remains highly challenging. Existing parameter-efficient adapters typically employ a single branch to model spatial and temporal cues jointly, which may fail to capture the fine-grained patterns of micro-gestures. To address this limitation, we propose a Spatial-Temporal Decoupled Adapter that decomposes video adaptation into independent temporal and spatial branches via lightweight depthwise convolutions. In addition, to alleviate the long-tailed class distribution inherent in the benchmark dataset, we introduce an Adaptive Soft Balanced Augmentation method, which dynamically allocates augmentation intensity based on class rarity and learning difficulty, without manual thresholds. Our method achieves an F1 score of 0.43808, ranking 1st in Track 2 of the 4th EI-MiGA-IJCAI Challenge.

微手势识别时空解耦自适应增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。