通过增强数据与集成学习,将强引力透镜检测误报率降低11倍。
Reducing false positives in strong lens detection through effective augmentation and ensemble learning
- 采用数据增强与集成学习提升训练数据多样性
- 误报率降至10^-4,真阳性保持88%以上
- 适合天文图像检测与大规模巡天项目应用
本研究探讨高质量训练数据对卷积神经网络(CNN)在强引力透镜检测中性能的影响。强调数据多样性和代表性的重要性,揭示样本分布变化对模型表现的影响。除数据质量外,实验验证了数据增强与集成学习在降低误报率的同时维持可接受的模型完整性。基于DenseNet和EfficientNet的实验达到最低误报率10^-4,测试集中成功识别超过88%的真实引力透镜,相比原始数据集误报率降低11倍,仅导致真阳性数量下降2.3%。该方法在KiDS数据集上验证有效,为欧几里得等未来巡天任务提供可借鉴方案。
原文摘要 · Abstract (English)
This research studies the impact of high-quality training datasets on the performance of Convolutional Neural Networks (CNNs) in detecting strong gravitational lenses. We stress the importance of data diversity and representativeness, demonstrating how variations in sample populations influence CNN performance. In addition to the quality of training data, our results highlight the effectiveness of various techniques, such as data augmentation and ensemble learning, in reducing false positives while maintaining model completeness at an acceptable level. This enhances the robustness of gravitational lens detection models and advancing capabilities in this field. Our experiments, employing variations of DenseNet and EfficientNet, achieved a best false positive rate (FP rate) of $10^{-4}$, while successfully identifying over 88 per cent of genuine gravitational lenses in the test dataset. This represents an 11-fold reduction in the FP rate compared to the original training dataset. Notably, this substantial enhancement in the FP rate is accompanied by only a 2.3 per cent decrease in the number of true positive samples. Validated on the KiDS dataset, our findings offer insights applicable to ongoing missions, like Euclid.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。