用自监督学习打造射电天文图像通用模型,提升形态分析能力。
STRADAViT: Towards a Foundational Model for Radio Astronomy through Self-Supervised Transfer
- 通过多源数据融合与射电天文特化视图生成训练视觉变换器。
- 在多个基准上优于ViT-MAE和DINOv2,RGZ DR1提升最显著。
- 适合需要跨望远镜、高可扩展性形态分析的研究者使用。
下一代射电天文巡天产生了数百万个分辨源,但跨异构望远镜和成像流程的鲁棒、可扩展形态分析仍具挑战。本文提出STRADAViT,一种基于自监督迁移的射电天文图像视觉变换器持续预训练框架。该框架整合多巡天数据筛选、射电天文感知的视图生成策略,以及以ViT-MAE初始化的编码器家族(支持注册标记)。包含重建仅、对比仅及两阶段分支。预训练数据集来自四个互补来源的射电图像切片。在三个形态分类基准上评估线性探查与微调性能,涵盖二分类与多分类场景。相较用于持续预训练的ViT-MAE初始化,最佳两阶段模型在所有线性探查设置中均提升宏观F1,在其中两个微调设置中表现更优,尤其在RGZ DR1上增益最大。相比DINOv2,提升具有选择性:在LoTSS DR2与RGZ DR1的线性探查,以及MiraBest与RGZ DR1的微调中均超越最强基线。针对性消融实验表明,适配方案不依赖初始点,且相同策略下,基于ViT-MAE的检查点因更低的令牌数量与下游成本被选为发布版本。结果表明,射电天文感知的视图生成与分阶段持续预训练能提供比现成视觉变换器更强的领域适应起点。
原文摘要 · Abstract (English)
Next-generation radio astronomy surveys are delivering millions of resolved sources, but robust and scalable morphology analysis remains difficult across heterogeneous telescopes and imaging pipelines. We present STRADAViT, a self-supervised Vision Transformer continued-pretraining framework for learning transferable encoders from radio astronomy imagery. The framework combines mixed-survey data curation, radio astronomy-aware training-view generation, and a ViT-MAE-initialized encoder family with optional register tokens. It supports reconstruction-only, contrastive-only, and two-stage branches. Our pretraining dataset comprises radio astronomy cutouts drawn from four complementary sources. We evaluate transfer with linear probing and fine-tuning on three morphology benchmarks spanning binary and multi-class settings. Relative to the ViT-MAE initialization used for continued pretraining, the best two-stage models improve Macro-F1 in all reported linear-probe settings and in two of three fine-tuning settings, with the largest gain on RGZ DR1. Relative to DINOv2, gains are selective rather than universal: the best two-stage models achieve higher mean Macro-F1 than the strongest DINOv2 baseline on LoTSS DR2 and RGZ DR1 under linear probing, and on MiraBest and RGZ DR1 under fine-tuning. A targeted DINOv2 initialization ablation further indicates that the adaptation recipe is not specific to the ViT-MAE starting point and that, under the same recipe. The ViT-MAE-based STRADAViT checkpoint is retained as the released checkpoint because it combines competitive transfer with substantially lower token count and downstream cost than the DINOv2-based alternative. These results indicate that radio astronomy-aware view generation and staged continued pretraining can provide a stronger domain-adapted starting point than off-the-shelf ViT checkpoints for radio astronomy transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。