用深度学习从海量射电数据中高效筛选脉冲星候选体
Pulsar Detection with Deep Learning
- 融合数组特征与图像诊断,构建端到端的脉冲星识别模型
- 最终系统准确率达94%,在少数类脉冲星上实现平衡的查全率与查准率
- 方法可推广至其他射电巡天项目,适合高通量数据实时筛查场景
脉冲星巡天每轮产生数百万候选体,人工甄别难以应对。本文构建了一套深度学习流水线,将阵列特征与图像诊断融合用于射电脉冲星候选体筛选。基于约500 GB的巨型米波射电望远镜(GMRT)数据,原始电压经SIGPROC转换为滤波器组,再通过PRESTO去弥散并折叠多个试选色散量,生成约3.2万个候选体。每个候选体生成四类诊断图:累加轮廓、时间-相位图、子带-相位图和色散量曲线,以数组与图像形式表示。基线堆叠模型(数组用ANN,图像用CNN,逻辑回归融合)准确率为68%。随后优化CNN架构与训练策略(正则化、学习率调度、最大范数约束),并通过针对性增强缓解类别不平衡,包括使用GAN生成少数类样本。改进后的CNN准确率达87%,最终的GAN+CNN系统在保留轻量化的同时,在独立测试集上达到94%准确率,且脉冲星类的精确率与召回率均衡。结果表明,融合数组与图像通道显著提升分类性能,适度生成增强可大幅提高少数类召回率。该方法具备巡天无关性,可扩展至未来高吞吐量观测设施。
原文摘要 · Abstract (English)
Pulsar surveys generate millions of candidates per run, overwhelming manual inspection. This thesis builds a deep learning pipeline for radio pulsar candidate selection that fuses array-derived features with image diagnostics. From approximately 500 GB of Giant Metrewave Radio Telescope (GMRT) data, raw voltages are converted to filterbanks (SIGPROC), then de-dispersed and folded across trial dispersion measures (PRESTO) to produce approximately 32,000 candidates. Each candidate yields four diagnostics--summed profile, time vs. phase, subbands vs. phase, and DM curve--represented as arrays and images. A baseline stacked model (ANNs for arrays + CNNs for images with logistic-regression fusion) reaches 68% accuracy. We then refine the CNN architecture and training (regularization, learning-rate scheduling, max-norm constraints) and mitigate class imbalance via targeted augmentation, including a GAN-based generator for the minority class. The enhanced CNN attains 87% accuracy; the final GAN+CNN system achieves 94% accuracy with balanced precision and recall on a held-out test set, while remaining lightweight enough for near--real-time triage. The results show that combining array and image channels improves separability over image-only approaches, and that modest generative augmentation substantially boosts minority (pulsar) recall. The methods are survey-agnostic and extensible to forthcoming high-throughput facilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。