用通用视觉语言模型零样本检测快速射电暴,效果接近专用模型。
Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector

- 用提示词控制通用模型,无需训练即可识别射电暴、干扰和噪声。
- 在模拟数据上准确率达94.05%,误报率远低于专用模型。
- 推理速度快于观测时间,适合实时处理大量天文数据。
快速射电暴(FRB)检测依赖专用深度学习模型,需大量特定数据训练且无法灵活调整。本文评估小型开源、本地运行的通用视觉语言模型(VLMs)在零样本提示模式下检测动态谱中FRB的可行性。在2000个模拟的L波段光谱的平衡二分类基准上,Gemma 4 E2B达到94.05%准确率,与专用检测器SwinYNet(92.85%)无统计差异,且在结构化射频干扰上的误报率显著更低(4.8% vs. 24.6%),纯噪声上为0;但SwinYNet ROC-AUC更高(1.0000 vs. 0.9520)。仅通过修改提示词,同一模型可实现三分类(FRB/RFI/噪声),准确率最高达86.0%,无误报,每2秒光谱处理仅需1.0–1.5秒,快于观测时长。未修改直接应用于1600个真实FAST-FREX观测,对真实干扰几乎完美排除(1000个负样本中仅2和5个误报),但仅恢复28.5%和27.0%的600条已知爆发,远低于SwinYNet报告的95.7%。按信号是否在图像中可辨识分层分析发现,性能瓶颈在于输入表示而非分类器——信号清晰时召回率达84–85%,信号不可见时降至1%。模拟爆发中位亮度近30倍于真实,亮度匹配时召回率相差仅数个百分点。
原文摘要 · Abstract (English)
Fast Radio Burst (FRB) detection increasingly relies on specialized deep learning models that require large task-specific training sets and cannot be redefined without retraining. We evaluate whether small, open-weight, locally run generalist Vision-Language Models (VLMs) can detect FRBs in dynamic spectra under a zero-shot, prompt-only regime. On a balanced binary benchmark of 2000 simulated L-band spectra, Gemma 4 E2B reaches an accuracy of 94.05\%, statistically indistinguishable from the specialized detector SwinYNet (92.85\%), with a far lower false-positive rate on structured RFI (4.8\% vs. 24.6\%) and none on pure noise, though SwinYNet ranks perfectly (ROC-AUC 1.0000 vs. 0.9520). Rewriting the prompt alone reconfigures the same models for three-class FRB/RFI/noise classification, reaching up to 86.0\% accuracy without a single false FRB while classifying each 2 s spectrum in 1.0--1.5 s, faster than the observation itself. Applied unchanged to the 1600 real FAST observations of FAST-FREX, they reject real interference almost perfectly (2 and 5 false positives in 1000 negatives) but recover only 28.5\% and 27.0\% of the 600 catalogued bursts, against 95.7\% reported for SwinYNet on the same files. Stratifying those bursts by the dispersed signal in the image shows the limit to be the input representation rather than the classifier, recall rising to 84--85\% where the sweep is unambiguous and collapsing to 1\% on the 13\% of positives carrying no detectable signal in a 2 s undedispersed full-band view. The simulated bursts are nearly 30 times brighter in median, and at matched brightness the recalls agree to within a few points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。