arXiv:2601.16210cs.CVcs.AI2026-01被引 5

多尺度视频离散化,让文字与视频对齐更准。

PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation

  • 用共享大二进制码本在多尺度上量化视频特征
  • 在10个基准上实现视频重建与生成新SOTA
  • 适合做跨模态视频理解与生成的开发者

离散视频变分自编码器(VAEs)支撑现代文生视频与视频理解系统,但现有分词器通常仅在单尺度学习视觉码本,词汇量有限且语言监督浅层,导致跨模态对齐差、零样本迁移能力弱。我们提出PyraTok,一种语言对齐的金字塔式分词器,在多个时空分辨率下学习语义结构化的离散潜在表示。PyraTok基于预训练视频VAE和新颖的语言对齐金字塔量化(LaPQ)模块,利用共享的大二进制码本在多个深度离散编码器特征,生成紧凑而丰富的视频标记序列。为紧密耦合视觉标记与语言,PyraTok联合优化多尺度文本引导量化与令牌层级上的全局自回归目标。在10个基准上,PyraTok实现视频重建新SOTA,持续提升文生视频质量,并在视频分割、时序动作定位和视频理解任务上达成新零样本SOTA,可稳定扩展至4K/8K分辨率。

原文摘要 · Abstract (English)

Discrete video VAEs underpin modern text-to-video generation and video understanding systems, yet existing tokenizers typically learn visual codebooks at a single scale with limited vocabularies and shallow language supervision, leading to poor cross-modal alignment and zero-shot transfer. We introduce PyraTok, a language-aligned pyramidal tokenizer that learns semantically structured discrete latents across multiple spatiotemporal resolutions. PyraTok builds on a pretrained video VAE and a novel Language aligned Pyramidal Quantization (LaPQ) module that discretizes encoder features at several depths using a shared large binary codebook, yielding compact yet expressive video token sequences. To tightly couple visual tokens with language, PyraTok jointly optimizes multi-scale text-guided quantization and a global autoregressive objective over the token hierarchy. Across ten benchmarks, PyraTok delivers state-of-the-art (SOTA) video reconstruction, consistently improves text-to-video quality, and sets new SOTA zero-shot performance on video segmentation, temporal action localization, and video understanding, scaling robustly to up to 4K/8K resolutions.

视频生成多尺度跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。