解决视频动作分割中长尾分布导致的分类偏差问题
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
- 引入学习状态感知的代价敏感损失,动态调整权重
- 在三个基准上实现帧级与段级性能显著提升
- 适合处理类别不均衡的视频动作识别任务
未剪辑流程类视频中的时间动作分割旨在密集标注帧为动作类别。这类视频天然存在长尾分布,动作频率和持续时间差异巨大。现有方法存在双层学习偏差:一是类别层面的偏差,由头类(高频)占优导致;二是转换层面的偏差,由常见转换模式主导。为此,我们提出一种约束优化方法,定义动作类别及其转换的学习状态,并融入优化过程。设计了一种新型代价敏感损失函数,基于动作与转换的学习状态自适应调整交叉熵权重。在三个具有挑战性的时序分割基准及多种框架上实验表明,该方法显著提升了每类动作的帧级与段级性能。
原文摘要 · Abstract (English)
Temporal action segmentation in untrimmed procedural videos aims to densely label frames into action classes. These videos inherently exhibit long-tailed distributions, where actions vary widely in frequency and duration. In temporal action segmentation approaches, we identified a bi-level learning bias. This bias encompasses (1) a class-level bias, stemming from class imbalance favoring head classes, and (2) a transition-level bias arising from variations in transitions, prioritizing commonly observed transitions. As a remedy, we introduce a constrained optimization problem to alleviate both biases. We define learning states for action classes and their associated transitions and integrate them into the optimization process. We propose a novel cost-sensitive loss function formulated as a weighted cross-entropy loss, with weights adaptively adjusted based on the learning state of actions and their transitions. Experiments on three challenging temporal segmentation benchmarks and various frameworks demonstrate the effectiveness of our approach, resulting in significant improvements in both per-class frame-wise and segment-wise performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。