让模拟星系图像更好匹配真实观测,提升星系分类准确率。
From Simulations to Surveys: Domain Adaptation for Galaxy Observations
- 用多种损失函数和配准策略对齐模拟与真实星系数据分布。
- 分类准确率从46%提升至87%,显著改善跨域性能。
- 适合天体物理、计算机视觉交叉研究者参考。
大规模测光巡天将拍摄数十亿个星系,但目前缺乏快速可靠的自动化方法来推断其形态、恒星质量及恒星形成率等物理属性。模拟生成的星系图像带有真实物理标签,但受点扩散函数、噪声、背景、选择偏差及标签先验差异等因素影响,存在显著领域偏移。本文提出初步的领域自适应流程:在模拟的TNG50星系上训练,评估在具有形态标签(椭圆/旋涡/不规则)的真实SDSS星系上的表现。采用三种主干网络(CNN、$E(2)$-可变形CNN、ResNet-18),结合焦点损失与有效类别数加权,并引入基于GeomLoss的特征级领域损失 $L_D$(包含熵正则化Sinkhorn最优传输、能量距离、高斯核MMD等)。结果表明,结合基于最优传输的“top_$k$软匹配”损失,聚焦于最差对齐的源-目标样本对,能进一步增强领域一致性。在欧氏距离、分阶段对齐权重和top-$k$匹配策略下,目标域准确率(宏F1)从无适配时的~46%(~30%)提升至~87%(~62.6%),领域判别AUC接近0.5,表明潜在空间高度混合。
原文摘要 · Abstract (English)
Large photometric surveys will image billions of galaxies, but we currently lack quick, reliable automated ways to infer their physical properties like morphology, stellar mass, and star formation rates. Simulations provide galaxy images with ground-truth physical labels, but domain shifts in PSF, noise, backgrounds, selection, and label priors degrade transfer to real surveys. We present a preliminary domain adaptation pipeline that trains on simulated TNG50 galaxies and evaluates on real SDSS galaxies with morphology labels (elliptical/spiral/irregular). We train three backbones (CNN, $E(2)$-steerable CNN, ResNet-18) with focal loss and effective-number class weighting, and a feature-level domain loss $L_D$ built from GeomLoss (entropic Sinkhorn OT, energy distance, Gaussian MMD, and related metrics). We show that a combination of these losses with an OT-based "top_$k$ soft matching" loss that focuses $L_D$ on the worst-matched source-target pairs can further enhance domain alignment. With Euclidean distance, scheduled alignment weights, and top-$k$ matching, target accuracy (macro F1) rises from $\sim$46% ($\sim$30%) at no adaptation to $\sim$87% ($\sim$62.6%), with a domain AUC near 0.5, indicating strong latent-space mixing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。