arXiv:2507.02862cs.CV2025-07被引 2

用参考帧提升视频生成的帧间一致性与细节保留能力。

RefTok: Reference-Based Tokenization for Video Generation

  • 以未量化参考帧为条件,动态编码解码视频帧序列。
  • 在4个数据集上平均提升36.7%的重建指标,压缩率相当或更高。
  • 适合需要高保真运动连续性和细节还原的视频生成任务。

有效处理时间冗余仍是视频模型学习的关键挑战。现有方法通常独立处理每组帧,难以捕捉视频中固有的时间依赖性和冗余性。为此,我们提出RefTok,一种基于参考帧的新型分词方法,能够捕捉复杂的时空动态和上下文信息。该方法将帧集编码和解码过程依赖于一个未量化的参考帧,解码后能保持运动连续性和物体外观的一致性。例如,即使头部移动也能保留面部细节,正确重建文本,保持小图案清晰,并维持手写字体可读性。在K600、UCF-101、BAIR Robot Pushing和DAVIS共4个视频数据集上,RefTok显著优于当前最先进分词器(Cosmos和MAGVIT),在相同或更高压缩率下,各项评估指标(PSNR、SSIM、LPIPS)平均提升36.7%。在BAIR Robot Pushing任务中,使用RefTok潜变量训练的生成模型,在所有生成指标上不仅超过MAGVIT-B,还超越参数量大4倍的MAGVIT-L,平均提升27.9%。

原文摘要 · Abstract (English)

Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectively capture the temporal dependencies and redundancies inherent in videos. To address this limitation, we introduce RefTok, a novel reference-based tokenization method capable of capturing complex temporal dynamics and contextual information. Our method encodes and decodes sets of frames conditioned on an unquantized reference frame. When decoded, RefTok preserves the continuity of motion and the appearance of objects across frames. For example, RefTok retains facial details despite head motion, reconstructs text correctly, preserves small patterns, and maintains the legibility of handwriting from the context. Across 4 video datasets (K600, UCF-101, BAIR Robot Pushing, and DAVIS), RefTok significantly outperforms current state-of-the-art tokenizers (Cosmos and MAGVIT) and improves all evaluated metrics (PSNR, SSIM, LPIPS) by an average of 36.7% at the same or higher compression ratios. When a video generation model is trained using RefTok's latents on the BAIR Robot Pushing task, the generations not only outperform MAGVIT-B but the larger MAGVIT-L, which has 4x more parameters, across all generation metrics by an average of 27.9%.

视频生成分词方法参考帧时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。