解决视频语言理解中数据量、多样性与质量的不可兼得难题
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
- 通过迭代优化视频标注,逐步提升数据质量
- 实现3%性能提升,多样性损失极小
- 适合大规模视频语言模型训练者使用
近期视频-语言理解得益于大规模预训练取得显著进展,但数据稀缺仍是主要挑战。本研究定量揭示了预训练数据集中数据量、多样性和质量之间的‘不可能三角’困境。现有方法试图通过合成标注改进高质量但低质量的ASR数据集,利用视频内容(帧、标签、语音转录等)来优化原始标注,但仍难以控制合成标注中的噪声,且在数据规模扩大时缺乏可扩展性。为此,本文提出Video DataFlywheel框架,通过迭代精炼视频标注并引入更优的噪声控制机制。首先用视频-语言模型生成合成标注构建优化数据集;然后在该数据集上预训练,并在人工标注样本上微调以获得更强模型;循环此过程实现持续改进。针对噪声控制,提出AdaTaiLr方法,对噪声分布假设更弱,理论保证更优,在大规模数据中表现更佳。结合迭代精炼与AdaTaiLr,显著提升可扩展性。大量实验表明,该框架优于现有数据精炼基线,实现3%性能提升,同时最小化多样性损失。所生成数据集在视频问答、文本-视频检索等任务中带来显著性能提升。
原文摘要 · Abstract (English)
Recently, video-language understanding has achieved great success through large-scale pre-training. However, data scarcity remains a prevailing challenge. This study quantitatively reveals an "impossible trinity" among data quantity, diversity, and quality in pre-training datasets. Recent efforts seek to refine large-scale, diverse ASR datasets compromised by low quality through synthetic annotations. These methods successfully leverage useful information in multimodal video content (frames, tags, ASR transcripts, etc.) to refine the original annotations. Nevertheless, they struggle to mitigate noise within synthetic annotations and lack scalability as the dataset size expands. To address these issues, we introduce the Video DataFlywheel framework, which iteratively refines video annotations with improved noise control methods. For iterative refinement, we first leverage a video-language model to generate synthetic annotations, resulting in a refined dataset. Then, we pre-train on it and fine-tune on human refinement examples for a stronger model. These processes are repeated for continuous improvement. For noise control, we present AdaTaiLr, a novel noise control method that requires weaker assumptions on noise distribution, thereby proving more effective in large datasets with theoretical guarantees. The combination of iterative refinement and AdaTaiLr can achieve better scalability in video-language understanding. Extensive experiments show that our framework outperforms existing data refinement baselines, delivering a 3% performance boost and improving dataset quality with minimal diversity loss. Furthermore, our refined dataset facilitates significant improvements in various video-language understanding tasks, including video question answering and text-video retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。