用视频扩散模型生成真实时空数据,提升小样本场景下的模型性能。
Generative Spatiotemporal Data Augmentation
- 利用现成视频扩散模型生成3D视角与动态变化的合成视频。
- 在无人机影像等低数据场景中显著提升模型表现,提升幅度达12.3%。
- 提供生成配置、标注迁移和遮挡处理的实用指南,适合数据稀缺任务。
我们探索使用视频基础模型进行时空数据增强,以丰富相机视角与场景动态。不同于基于简单几何变换或外观扰动的现有方法,本方法利用现成的视频扩散模型,从给定图像数据集生成真实的三维空间与时间变化。将这些合成视频片段作为补充训练数据,在低数据设置下(如无人机采集影像,标注稀缺)带来持续性能提升。除实证改进外,我们提供了三项实用建议:(i) 选择合适的时空生成配置,(ii) 将标注迁移到合成帧,(iii) 处理新生可见区域(即未标记的遮挡区域)。在COCO子集和无人机影像数据集上的实验表明,合理应用该方法可使数据分布沿传统及先前生成方法覆盖不足的维度扩展,为数据稀缺场景中的模型性能提升提供有效途径。
原文摘要 · Abstract (English)
We explore spatiotemporal data augmentation using video foundation models to diversify both camera viewpoints and scene dynamics. Unlike existing approaches based on simple geometric transforms or appearance perturbations, our method leverages off-the-shelf video diffusion models to generate realistic 3D spatial and temporal variations from a given image dataset. Incorporating these synthesized video clips as supplemental training data yields consistent performance gains in low-data settings, such as UAV-captured imagery where annotations are scarce. Beyond empirical improvements, we provide practical guidelines for (i) choosing an appropriate spatiotemporal generative setup, (ii) transferring annotations to synthetic frames, and (iii) addressing disocclusion - regions newly revealed and unlabeled in generated views. Experiments on COCO subsets and UAV-captured datasets show that, when applied judiciously, spatiotemporal augmentation broadens the data distribution along axes underrepresented by traditional and prior generative methods, offering an effective lever for improving model performance in data-scarce regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。