arXiv:2512.11145cs.LGcs.AI2025-12被引 1

用自监督方法提升科学集合数据的可视化与可解释性。

SENSE: Self-Supervised Neural Embeddings for Spatial Ensembles

  • 结合聚类与对比损失,优化自动编码器以更好捕捉数据结构。
  • 在土壤通道与液滴冲击数据集上,新模型显著提升聚类效果。
  • 适合处理高维科学数据的可视化分析,尤其关注结构可解释性。

高维复杂科学集合数据的分析与可视化面临巨大挑战。尽管降维技术和自动编码器是提取特征的强大工具,但对高维数据仍常表现不佳。本文提出一种增强型自动编码器框架,引入基于软轮廓系数的聚类损失与对比损失,以提升集合数据的可视化与可解释性。首先,使用EfficientNetV2为未标记数据生成伪标签;通过联合优化重建、聚类和对比目标,使相似数据点在潜在空间中聚集,不同簇间保持分离。随后,利用UMAP对潜在表示进行二维投影,并以轮廓系数评估效果。在两个科学集合数据集(基于马尔可夫链蒙特卡洛的土壤通道结构,以及液滴-薄膜撞击动力学)上测试多种自动编码器,结果表明加入聚类或对比损失的模型相较基线方法略有提升。

原文摘要 · Abstract (English)

Analyzing and visualizing scientific ensemble datasets with high dimensionality and complexity poses significant challenges. Dimensionality reduction techniques and autoencoders are powerful tools for extracting features, but they often struggle with such high-dimensional data. This paper presents an enhanced autoencoder framework that incorporates a clustering loss, based on the soft silhouette score, alongside a contrastive loss to improve the visualization and interpretability of ensemble datasets. First, EfficientNetV2 is used to generate pseudo-labels for the unlabeled portions of the scientific ensemble datasets. By jointly optimizing the reconstruction, clustering, and contrastive objectives, our method encourages similar data points to group together while separating distinct clusters in the latent space. UMAP is subsequently applied to this latent representation to produce 2D projections, which are evaluated using the silhouette score. Multiple types of autoencoders are evaluated and compared based on their ability to extract meaningful features. Experiments on two scientific ensemble datasets - channel structures in soil derived from Markov chain Monte Carlo, and droplet-on-film impact dynamics - show that models incorporating clustering or contrastive loss marginally outperform the baseline approaches.

自监督学习科学可视化聚类自动编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。