arXiv:2412.10443cs.CVcs.AI2024-12ICCV被引 14

提出新型视频分词器SweetTok,实现高效压缩与高保真重建。

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization

  • 分离空间与时间查询,通过解耦自编码器压缩视频数据。
  • 在UCF-101上重建误差降低42.8%,生成质量提升15.1%。
  • 生成带语义的紧凑令牌,适合少样本识别等下游任务。

本文提出一种新型视频分词器SweetTok,以克服现有方法在紧凑且高效离散化方面的局限。与以往直接对局部视觉块或自适应查询进行分词不同,SweetTok采用解耦框架,通过独立的空间与时间查询,利用解耦查询自编码器(DQAE)压缩视觉输入,从而在显著减少视频令牌数量的同时,保持更优的保真度。此外,设计了专为时空压缩优化的运动增强语言码本(MLC),以应对外观与运动信息在语义表征上的差异。实验表明,SweetTok在UCF-101数据集上使重建指标rFVD降低42.8%;在下游视频生成任务中,生成指标gFVD提升15.1%。同时,压缩后的解耦令牌蕴含语义信息,可支持基于大语言模型的少样本识别应用。

原文摘要 · Abstract (English)

This paper presents the \textbf{S}emantic-a\textbf{W}ar\textbf{E} spatial-t\textbf{E}mporal \textbf{T}okenizer (SweetTok), a novel video tokenizer to overcome the limitations in current video tokenization methods for compacted yet effective discretization. Unlike previous approaches that process flattened local visual patches via direct discretization or adaptive query tokenization, SweetTok proposes a decoupling framework, compressing visual inputs through distinct spatial and temporal queries via \textbf{D}ecoupled \textbf{Q}uery \textbf{A}uto\textbf{E}ncoder (DQAE). This design allows SweetTok to efficiently compress video token count while achieving superior fidelity by capturing essential information across spatial and temporal dimensions. Furthermore, we design a \textbf{M}otion-enhanced \textbf{L}anguage \textbf{C}odebook (MLC) tailored for spatial and temporal compression to address the differences in semantic representation between appearance and motion information. SweetTok significantly improves video reconstruction results by \textbf{42.8\%} w.r.t rFVD on UCF-101 dataset. With a better token compression strategy, it also boosts downstream video generation results by \textbf{15.1\%} w.r.t gFVD. Additionally, the compressed decoupled tokens are imbued with semantic information, enabling few-shot recognition capabilities powered by LLMs in downstream applications.

视频分词压缩生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。