分离传感器噪声与真实信号,提升多源天文图像分析精度
Learning What's Real: Disentangling Signal and Measurement Artifacts in Multi-Sensor Data, with Applications to Astrophysics
- 用双重编码器和反事实生成目标,从多传感器数据中解耦真实信号与测量干扰
- 在DESILegacy与HSC两套巡天数据上验证,实现无仪器依赖的图像相似性搜索
- 适用于天文、医学等多模态科学数据的自监督预训练,特别适合跨设备分析
真实世界的数据总是包含两部分:感兴趣的物理过程本身产生的信号,以及由传感器或仪器引起的测量相关伪影。后者作为混杂因素,限制了对物理本质信息的提取,并使异构或多元仪器观测的融合变得复杂。本文提出一种深度学习框架,利用重叠观测、双编码器结构及反事实生成目标,解耦这些变异因素。所得表征可明确区分内在信号与传感器特异性失真和噪声,可用于反事实视图生成、不受测量畸变影响的参数推断,以及仪器无关的相似性搜索。我们在来自DESI Legacy Imaging Survey(Legacy)和Hyper Suprime-Cam(HSC)Survey的天体星系图像上进行了验证,展示了该方法在典型多仪器场景下的有效性。该框架为科学与多模态自监督预训练提供通用范式:从同一物理系统的重叠观测中构造训练样本,将传感器或模态特异性效应视为增强,通过反事实生成学习不变表征。
原文摘要 · Abstract (English)
Data collected from the physical world is always a combination of multiple sources: an underlying signal from the physical process of interest and a signal from measurement-dependent artifacts from the sensor or instrument. This secondary signal acts as a confounding factor, limiting our ability to extract information about the physics underlying the phenomena we observe. Furthermore, it complicates the combination of observations in heterogeneous or multi-instrument settings. We propose a deep learning framework that leverages overlapping observations, a dual-encoder architecture, and a counterfactual generation objective to disentangle these factors of variation. The resulting representations explicitly separate intrinsic signals from sensor-specific distortions and noise, and can be used for counterfactual view generation, parameter inference unconfounded by measurement distortions, and instrument-independent similarity search. We demonstrate the effectiveness of our approach on astrophysical galaxy images from the DESI Legacy Imaging Survey (Legacy) and the Hyper Suprime-Cam (HSC) Survey as a representative multi-instrument setting. This framework provides a general recipe for scientific and multi-modal self-supervised pretraining: construct training pairs from overlapping observations of the same physical system, treat sensor- or modality-specific effects as augmentations, and learn invariant representations through counterfactual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。