首个统一视觉分词器,跨图像视频3D实现高保真重建与语义理解。
AToken: A Unified Tokenizer for Vision

- 用4D旋转位置编码的纯Transformer架构处理多模态视觉输入。
- 图像重建rFID达0.21,视频检索准确率40.2%,3D重建PSNR达28.28。
- 支持生成与理解双重任务,适合多模态大模型和跨模态应用。
我们提出AToken,首个统一视觉分词器,可在图像、视频和3D资产上同时实现高保真重建与语义理解。不同于仅专注重建或理解的单模态分词器,AToken将多种视觉输入编码至共享4D隐空间,统一任务与模态。其采用纯Transformer架构与4D旋转位置嵌入,可处理任意分辨率与时间长度的输入。为保障训练稳定,引入无对抗训练目标,结合感知损失与格拉姆矩阵损失,达到当前最优重建质量。通过渐进式训练流程,AToken从单图、视频到3D逐步扩展,支持连续与离散隐码。在图像上实现0.21 rFID与82.2% ImageNet准确率,视频上3.01 rFVD与40.2% MSRVTT检索准确率,3D上28.28 PSNR与90.9%分类准确率。下游应用中,支持图像生成(连续/离散码)、文本到视频、图像到3D合成及多模态大模型理解,各基准表现均具竞争力。该工作为下一代统一视觉表征的多模态系统提供了新范式。
原文摘要 · Abstract (English)
We present AToken, the first unified visual tokenizer that achieves both high-fidelity reconstruction and semantic understanding across images, videos, and 3D assets. Unlike existing tokenizers that specialize in either reconstruction or understanding for single modalities, AToken encodes these diverse visual inputs into a shared 4D latent space, unifying both tasks and modalities in a single framework. Specifically, we introduce a pure transformer architecture with 4D rotary position embeddings to process visual inputs of arbitrary resolutions and temporal durations. To ensure stable training, we introduce an adversarial-free training objective that combines perceptual and Gram matrix losses, achieving state-of-the-art reconstruction quality. By employing a progressive training curriculum, AToken gradually expands from single images, videos, and 3D, and supports both continuous and discrete latent tokens. AToken achieves 0.21 rFID with 82.2% ImageNet accuracy for images, 3.01 rFVD with 40.2% MSRVTT retrieval for videos, and 28.28 PSNR with 90.9% classification accuracy for 3D.. In downstream applications, AToken enables both visual generation tasks (e.g., image generation with continuous and discrete tokens, text-to-video generation, image-to-3D synthesis) and understanding tasks (e.g., multimodal LLMs), achieving competitive performance across all benchmarks. These results shed light on the next-generation multimodal AI systems built upon unified visual tokenization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。