VidTok高效压缩视频为紧凑令牌,支持连续与离散两种方式。
VidTok: A Versatile and Open-Source Video Tokenizer
- 采用卷积与上下采样结构,提升视频编码效率。
- 使用FSQ避免码本坍塌,训练更稳定,性能优于现有方法。
- 适合视频生成与理解研究者,开源可用。
将视频内容编码为紧凑的潜在令牌已成为视频生成与理解的关键步骤,以应对像素级表示中的固有冗余。随着以视频为中心的研究日益重要,对高性能、开源视频令牌化工具的需求不断增长。我们提出VidTok,一种在连续和离散令牌化方面均达到顶尖性能的多功能视频令牌化器。VidTok在现有方法基础上实现多项改进:1)引入卷积层与上下采样模块的模型架构;2)为解决传统向量量化(VQ)中常见的训练不稳定与码本坍塌问题,采用有限标量量化(FSQ)进行离散视频令牌化;3)优化训练策略,包括两阶段训练流程及降低帧率的训练方式。通过整合这些改进,VidTok在标准评估设置下显著优于现有方法,在多个指标(如PSNR、SSIM、LPIPS、FVD)上表现更优。
原文摘要 · Abstract (English)
Encoding video content into compact latent tokens has become a fundamental step in video generation and understanding, driven by the need to address the inherent redundancy in pixel-level representations. Consequently, there is a growing demand for high-performance, open-source video tokenizers as video-centric research gains prominence. We introduce VidTok, a versatile video tokenizer that delivers state-of-the-art performance in both continuous and discrete tokenizations. VidTok incorporates several key advancements over existing approaches: 1) model architecture such as convolutional layers and up/downsampling modules; 2) to address the training instability and codebook collapse commonly associated with conventional Vector Quantization (VQ), we integrate Finite Scalar Quantization (FSQ) into discrete video tokenization; 3) improved training strategies, including a two-stage training process and the use of reduced frame rates. By integrating these advancements, VidTok achieves substantial improvements over existing methods, demonstrating superior performance across multiple metrics, including PSNR, SSIM, LPIPS, and FVD, under standardized evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。