arXiv:2508.04181cs.CV2025-08

探究大模型ViT-22B在本地训练中的稳定性与图像生成潜力

Deeper Inside Deep ViT

  • 在本地环境下分析ViT-22B训练行为并提出稳定化改进方法
  • 同参数量下ViT-22B性能优于普通ViT,且可支持图像生成任务
  • 首次尝试用ViT-22B进行图像生成,验证其结构适配性

已有研究尝试构建类似大语言模型的大型视觉模型,如ViT-22B。尽管相关工作提供了大量分析与洞见,但其实际应用价值仍不明确。本文在本地环境中考察该模型的训练表现,揭示训练不稳定性问题,并提出相应改进措施以增强稳定性。实验表明,从头训练的ViT-22B在相同参数规模下性能优于标准ViT。此外,本文首次探索了将ViT-22B应用于图像生成任务,提出基于ViT的生成架构,并对比ViT与ViT-22B在生成任务中的适用性,为大规模视觉模型的多功能扩展提供新思路。

原文摘要 · Abstract (English)

There have been attempts to create large-scale structures in vision models similar to LLM, such as ViT-22B. While this research has provided numerous analyses and insights, our understanding of its practical utility remains incomplete. Therefore, we examine how this model structure reacts and train in a local environment. We also highlight the instability in training and make some model modifications to stabilize it. The ViT-22B model, trained from scratch, overall outperformed ViT in terms of performance under the same parameter size. Additionally, we venture into the task of image generation, which has not been attempted in ViT-22B. We propose an image generation architecture using ViT and investigate which between ViT and ViT-22B is a more suitable structure for image generation.

视觉模型大模型图像生成ViT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。