通过多样化数据增强与领域自适应,缓解视频时序定位中的时间偏差问题。
Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding
- 设计多维度增强策略,生成不同长度和目标位置的视频以打破时间分布偏倚。
- 在Charades-CD和ActivityNet-CD上实现最新性能,提升模型泛化能力。
- 适合关注视频理解中公平性与鲁棒性的研究者和开发者。
视频时序定位(TSGV)因公开数据集存在显著的时间偏倚而面临挑战,这源于目标片段在时间分布上的不均衡。现有方法通过生成视频并强制改变目标片段的位置来缓解此问题,但由于数据集视频长度变化小,仅调整位置导致模型对长度多变的视频泛化能力差。本文提出一种融合多样化数据增强与领域判别器的新训练框架:增强策略生成具有不同长度和目标位置的视频以丰富时间分布;为应对增强后特征分布差异带来的噪声,引入领域自适应辅助任务以减少原始与增强视频间的特征差距;同时鼓励模型对相同文本查询但不同位置的视频产生差异化预测,促进去偏训练。在Charades-CD和ActivityNet-CD数据集上的实验表明,该方法在多种定位结构下均表现出优异的有效性与泛化能力,达到当前最优效果。
原文摘要 · Abstract (English)
Temporal sentence grounding in videos (TSGV) faces challenges due to public TSGV datasets containing significant temporal biases, which are attributed to the uneven temporal distributions of target moments. Existing methods generate augmented videos, where target moments are forced to have varying temporal locations. However, since the video lengths of the given datasets have small variations, only changing the temporal locations results in poor generalization ability in videos with varying lengths. In this paper, we propose a novel training framework complemented by diversified data augmentation and a domain discriminator. The data augmentation generates videos with various lengths and target moment locations to diversify temporal distributions. However, augmented videos inevitably exhibit distinct feature distributions which may introduce noise. To address this, we design a domain adaptation auxiliary task to diminish feature discrepancies between original and augmented videos. We also encourage the model to produce distinct predictions for videos with the same text queries but different moment locations to promote debiased training. Experiments on Charades-CD and ActivityNet-CD datasets demonstrate the effectiveness and generalization abilities of our method in multiple grounding structures, achieving state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。