arXiv:2602.16918cs.CVcs.AI2026-02

用150亿图文对训练大模型,提升跨模态理解与鲁棒性。

Xray-Visual Models: Scaling Vision models on Industry Scale Data

  • 三阶段训练:自监督、半监督标签分类、对比学习联合优化
  • 在ImageNet、Kinetics等9个基准上达领先水平,视频任务提升6.2%
  • 适合工业级多模态应用,尤其适配真实复杂环境的检索需求

我们提出Xray-Visual,一种统一的视觉模型架构,用于在大规模社交媒体数据上进行图像与视频理解。模型基于来自Facebook和Instagram的超过150亿条精心筛选的图文对以及100亿条视频-标签对,采用稳健的数据清洗流程,包含平衡策略与噪声抑制,以最大化语义多样性并最小化标签噪声。引入三阶段训练流程:自监督MAE、半监督标签分类与类似CLIP的对比学习,联合优化图像与视频模态。架构基于带高效令牌重组(EViT)的Vision Transformer,提升计算效率。大量实验表明,Xray-Visual在多个基准上达到当前最优表现,包括ImageNet图像分类、Kinetics与HMDB51视频理解,以及MSCOCO跨模态检索。模型对域偏移和对抗扰动具有强鲁棒性。进一步证明,使用大语言模型作为文本编码器(LLM2CLIP)可显著提升检索性能与泛化能力,尤其在真实环境中表现突出。Xray-Visual为可扩展的多模态视觉模型树立新标准,同时保持高精度与高效率。

原文摘要 · Abstract (English)

We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 billion curated image-text pairs and 10 billion video-hashtag pairs from Facebook and Instagram, employing robust data curation pipelines that incorporate balancing and noise suppression strategies to maximize semantic diversity while minimizing label noise. We introduce a three-stage training pipeline that combines self-supervised MAE, semi-supervised hashtag classification, and CLIP-style contrastive learning to jointly optimize image and video modalities. Our architecture builds on a Vision Transformer backbone enhanced with efficient token reorganization (EViT) for improved computational efficiency. Extensive experiments demonstrate that Xray-Visual achieves state-of-the-art performance across diverse benchmarks, including ImageNet for image classification, Kinetics and HMDB51 for video understanding, and MSCOCO for cross-modal retrieval. The model exhibits strong robustness to domain shift and adversarial perturbations. We further demonstrate that integrating large language models as text encoders (LLM2CLIP) significantly enhances retrieval performance and generalization capabilities, particularly in real-world environments. Xray-Visual establishes new benchmarks for scalable, multimodal vision models, while maintaining superior accuracy and computational efficiency.

视觉模型多模态工业级数据Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。