用3D分形视频预训练动作识别模型,提升效率与效果
How to Sample High Quality 3D Fractals for Action Recognition Pre-Training?
- 通过3DIFS生成动态分形视频,构建无标注成本的预训练数据
- 提出定向智能过滤法,速度提升约100倍,下游任务性能更优
- 适合需要高效合成数据的视觉预训练研究者
在深度学习领域,合成数据正成为耗时且昂贵的真实标注数据的有力替代。其中,公式驱动监督学习(FDSL)可通过公式(如分形或轮廓)生成无限数量的完美标注数据,避免人工标注、隐私及伦理问题。本文利用3D迭代函数系统(3D IFS)生成3D分形,并将其时间变换为视频,用于动作识别模型的预训练。发现传统分形生成方法速度慢且易产生退化结果。因此系统性探索替代生成方式,发现过于严格的生成策略虽美观,却损害下游任务表现。为此提出新颖的定向智能过滤(Targeted Smart Filtering)方法,兼顾生成速度与分形多样性,实现约100倍的采样速度提升,并在下游动作识别任务中优于其他3D分形过滤方法。
原文摘要 · Abstract (English)
Synthetic datasets are being recognized in the deep learning realm as a valuable alternative to exhaustively labeled real data. One such synthetic data generation method is Formula Driven Supervised Learning (FDSL), which can provide an infinite number of perfectly labeled data through a formula driven approach, such as fractals or contours. FDSL does not have common drawbacks like manual labor, privacy and other ethical concerns. In this work we generate 3D fractals using 3D Iterated Function Systems (IFS) for pre-training an action recognition model. The fractals are temporally transformed to form a video that is used as a pre-training dataset for downstream task of action recognition. We find that standard methods of generating fractals are slow and produce degenerate 3D fractals. Therefore, we systematically explore alternative ways of generating fractals and finds that overly-restrictive approaches, while generating aesthetically pleasing fractals, are detrimental for downstream task performance. We propose a novel method, Targeted Smart Filtering, to address both the generation speed and fractal diversity issue. The method reports roughly 100 times faster sampling speed and achieves superior downstream performance against other 3D fractal filtering methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。