用生成视频平衡长尾数据,显著提升罕见动作识别效果。
Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

- 用文本生成视频补足稀有动作样本,构建均衡训练集。
- 在UCF-LT和K100-LT上分别提升5.1%和7.0%准确率。
- 少量合成数据即可达成近80%性能增益,计算成本仅27%。
针对视频动作识别中的长尾数据问题,本文提出Gen2Balance方法:利用文本到视频生成模型,基于动作特征和真实样本生成多样化视频,将不平衡数据集转化为真实与生成视频的均衡组合。为有效学习此类数据,采用两阶段训练策略以缓解域偏移,显著提升性能。在标准基准的长尾版本UCF-LT和K100-LT(聚焦时间挑战性动作)上,相较于最强基线,分别提升5.1%和7.0%准确率。在RareAct数据集的罕见动作(如切键盘)上,准确率提升达31.9%。通过调整合成数据量,发现仅27%计算成本即可获得79%性能增益,证明该方法具备良好的实用性与可扩展性。
原文摘要 · Abstract (English)
We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on diverse text prompts grounded in action profiles and training exemplars. Our approach, called Gen2Balance, converts an imbalanced training set into a balanced combination of real and generated video clips. To effectively learn from such data, we employ a two-stage training strategy that mitigates domain shift and yields significant improvements. We evaluate on long-tailed versions of standard benchmarks: UCF-101 (UCF-LT) and a 100-class subset of Kinetics (K100-LT) selected to prioritise temporally challenging actions. Gen2Balance improves accuracy over the strongest baselines for long-tailed learning by 5.1% and 7.0% on the respective datasets. On rare actions from the RareAct dataset (e.g., cut keyboard), Gen2Balance improves accuracy by 31.9%, demonstrating effectiveness for scarce actions. By varying the amount of synthetic data added, we show that partial balancing already achieves 79% of the performance gains at 27% of the compute cost on K100-LT, highlighting the practical scalability of Gen2Balance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。