arXiv:2606.30811cs.CVcs.MM2026-06被引 1

统一音视频分词器,让声音与画面共享编码空间。

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

论文配图:AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
图 1 · 摘自论文原文
  • 用共享编码器+专属查询,将音视频转为统一一维向量。
  • 训练时分阶段重建,解决音视频信息不均衡问题。
  • 适合构建统一的音视频生成大模型,提升同步精度。

音视频生成近年来备受关注,旨在合成高质量、声画高度同步且语义对齐的视听内容。现有方法多采用双分支设计,各模态独立分词与生成,忽略表征差距,且需大量计算资源进行训练。受一维视觉分词启发,我们提出 extbf{AVTok},一种面向整体音视频生成的新型统一分词器。AVTok 采用双流变压器架构,共享编码器-解码器并引入模态特异性可学习查询,高效将音视频对编码为紧凑的一维潜在表示,并使用统一码本。为应对阻碍音视频对齐信息利用的异质性信息不平衡问题,我们设计分层训练策略,逐步实现各模态的重建能力。大量实验表明,AVTok 在音视频重建以及下游任务(音频到视频、视频到音频、类别条件联合生成)中均表现优异。AVTok 为联合音视频分词挑战提供新路径,并有望推动统一大规模多模态音视频生成模型的发展。

原文摘要 · Abstract (English)

Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained synchronization and semantic alignment between the auditory and visual components. The preceding methods predominantly adopt a dual-branch design with separate tokenization and generation modules per modality, neglecting the representation gap while necessitating intensive computational resources for proper training. Inspired by recent advancements in one-dimensional visual tokenization, we present \textbf{AVTok}, a novel unified tokenizer designated for holistic audio-video generation. AVTok features a dual-stream transformer-based architecture with shared encoder-decoder and modal-specific learnable queries to efficiently and effectively encode an audio-video pair into a compact one-dimensional latent representation with a unified codebook. To cope with the heterogeneous information imbalance that hinders AVTok from exploiting aligned audio-visual information, we devise a hierarchical training strategy to progressively realize reconstruction capabilities for each modality. Extensive experiments demonstrate that AVTok excels both in audio-video reconstruction and when integrated into downstream pipelines for audio-to-video, video-to-audio, and class-conditional joint audio-video generation. AVTok paves the way for the challenge of joint audio-video tokenization and provides a potential direction to build unified large multimodal models for audio-video generation.

音视频生成统一分词多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。