arXiv:2505.18142cs.CVcs.DB2025-05被引 8

提出TokBench基准,评估视觉分词器对文字和人脸的重建能力。

TokBench: Evaluating Your Visual Tokenizer before Visual Generation

  • 用文字和人脸图像构建评估基准,聚焦细粒度特征保留。
  • 仅需2GB内存和4分钟完成测试,效率高且准确。
  • 发现当前分词器在小尺度下仍难保真,适合生成模型开发者使用。

本文揭示了视觉分词器与变分自编码器(VAE)在保留细粒度特征方面的局限性,并提出一个评估基准,用于衡量其在文本和人脸这两种挑战性视觉内容上的重建性能。尽管这些技术通过压缩或量化图像表示显著提升了视觉生成与多模态建模效率,但图像压缩带来的信息损失从根本上限制了生成质量上限。为评估该上限,研究聚焦于文本与人脸特征,因其具有小尺度、密集纹理、易崩溃及对人眼敏感等特点。研究从现有数据集收集并整理清晰的文本与人脸图像,采用成熟的OCR与人脸识别模型进行评估,确保准确性的同时实现极轻量级流程,仅需2GB内存和4分钟即可完成。基于此基准,分析不同分词器与VAE在多种尺度下的重建表现,结果表明现代视觉分词器在小尺度下仍难以保持细粒度特征。研究进一步将框架扩展至视频领域,全面评估视频分词器。此外,证明传统指标无法准确反映文本与人脸的重建效果,而新提出的指标可有效补充。

原文摘要 · Abstract (English)

In this work, we reveal the limitations of visual tokenizers and VAEs in preserving fine-grained features, and propose a benchmark to evaluate reconstruction performance for two challenging visual contents: text and face. Visual tokenizers and VAEs have significantly advanced visual generation and multimodal modeling by providing more efficient compressed or quantized image representations. However, while helping production models reduce computational burdens, the information loss from image compression fundamentally limits the upper bound of visual generation quality. To evaluate this upper bound, we focus on assessing reconstructed text and facial features since they typically: 1) exist at smaller scales, 2) contain dense and rich textures, 3) are prone to collapse, and 4) are highly sensitive to human vision. We first collect and curate a diverse set of clear text and face images from existing datasets. Unlike approaches using VLM models, we employ established OCR and face recognition models for evaluation, ensuring accuracy while maintaining an exceptionally lightweight assessment process <span style="font-weight: bold; color: rgb(214, 21, 21);">requiring just 2GB memory and 4 minutes</span> to complete. Using our benchmark, we analyze text and face reconstruction quality across various scales for different image tokenizers and VAEs. Our results show modern visual tokenizers still struggle to preserve fine-grained features, especially at smaller scales. We further extend this evaluation framework to video, conducting comprehensive analysis of video tokenizers. Additionally, we demonstrate that traditional metrics fail to accurately reflect reconstruction performance for faces and text, while our proposed metrics serve as an effective complement.

视觉分词器图像重建评测基准文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。