arXiv:2601.17237cs.CV2026-01被引 9

C-RADIOv4通过多教师蒸馏,提升视觉模型性能与灵活性。

C-RADIOv4 (Tech Report)

  • 采用多教师蒸馏融合SigLIP2、DINOv3、SAM3优势
  • 在相同算力下,下游任务表现显著提升,支持任意分辨率
  • 新增ViTDet选项,高分辨率下效率大幅优化,开源可商用

本文介绍C-RADIO系列最新版本C-RADIOv4,基于AM-RADIO/RADIOv2.5设计,通过多教师蒸馏构建统一学生模型,保留并增强多个教师的独特能力。新版本释放- SO400M(412M参数)和-H(631M)两种变体,训练使用更新的教师模型:SigLIP2、DINOv3和SAM3。相比前代,在核心指标上实现显著提升,并因模仿SAM3引入新能力。此外,进一步强化任意分辨率支持,重新加入ViTDet选项,大幅提高高分辨率推理效率,且采用宽松许可协议,便于广泛应用。

原文摘要 · Abstract (English)

By leveraging multi-teacher distillation, agglomerative vision backbones provide a unified student model that retains and improves the distinct capabilities of multiple teachers. In this tech report, we describe the most recent release of the C-RADIO family of models, C-RADIOv4, which builds upon AM-RADIO/RADIOv2.5 in design, offering strong improvements on key downstream tasks at the same computational complexity. We release -SO400M (412M params), and -H (631M) model variants, both trained with an updated set of teachers: SigLIP2, DINOv3, and SAM3. In addition to improvements on core metrics and new capabilities from imitating SAM3, the C-RADIOv4 model family further improves any-resolution support, brings back the ViTDet option for drastically enhanced efficiency at high-resolution, and comes with a permissive license.

视觉模型蒸馏多教师高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。