用无监督方法从ICESat-2数据中提取海冰特征嵌入,减少标注依赖。
Exploring the Potential of Latent Embeddings for Sea Ice Characterization using ICESat-2 Data
- 用LSTM和CNN自编码器重建海冰高程序列,生成低维嵌入。
- 嵌入后聚类更紧凑,保留整体结构,提升特征可区分性。
- 适合希望减少人工标注的极地遥感研究者使用。
ICESat-2 提供高分辨率海冰高度测量数据。现有研究多基于机器学习对海冰表面类型进行分类,但依赖大量人工标注,需交叉比对轨道数据与重叠光学影像,耗时费力。由于过境轨迹和大气条件差异,两者重合率较低。为解决此问题,本研究探索无监督自编码器在未标记数据上的应用,以提取潜在嵌入。基于长短期记忆网络(LSTM)和卷积神经网络(CNN)构建自编码器模型,重建ICESat-2地形序列并生成嵌入表示。进一步采用均匀流形近似与投影(UMAP)进行降维可视化。结果表明,自编码器生成的嵌入保留了原始数据的整体结构,且聚类更为紧凑,显示其有望显著降低标注样本需求。
原文摘要 · Abstract (English)
The Ice, Cloud, and Elevation Satellite-2 (ICESat-2) provides high-resolution measurements of sea ice height. Recent studies have developed machine learning methods on ICESat-2 data, primarily focusing on surface type classification. However, the heavy reliance on manually collected labels requires significant time and effort for supervised learning, as it involves cross-referencing track measurements with overlapping background optical imagery. Additionally, the coincidence of ICESat-2 tracks with background images is relatively rare due to the different overpass patterns and atmospheric conditions. To address these limitations, this study explores the potential of unsupervised autoencoder on unlabeled data to derive latent embeddings. We develop autoencoder models based on Long Short-Term Memory (LSTM) and Convolutional Neural Networks (CNN) to reconstruct topographic sequences from ICESat-2 and derive embeddings. We then apply Uniform Manifold Approximation and Projection (UMAP) to reduce dimensions and visualize the embeddings. Our results show that embeddings from autoencoders preserve the overall structure but generate relatively more compact clusters compared to the original ICESat-2 data, indicating the potential of embeddings to lessen the number of required labels samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。