用变分自编码器解析百万级X射线散射数据,实现结构演变的可视化与生成。
Unlocking Latent Dimensions: Exploring Representations of Large-Scale X-ray Scattering Data using Variational Autoencoders

- 基于注意力机制的卷积变分自编码器,从150万张图像中学习低维表示。
- 模型在两个同步辐射设施上成功组织时间分辨实验数据,呈现可解释的结构演化轨迹。
- 相比通用视觉模型,领域专用训练使潜空间更具物理意义,适合材料科学实时分析。
科学用户设施产生的X射线散射数据速度远超传统流程处理能力。本文针对离线数据探索和在线实时分析两种场景,利用150万张X射线散射图像训练了一个领域特定的注意力式卷积变分自编码器(C-VAE),以学习捕捉不同实验条件下结构变化的低维表示。所学潜空间呈现出有序聚类和反映实验进程的平滑轨迹,并支持跨多种结构状态的可控合成散射图像生成。该模型无需重新训练即可将两个同步辐射设施的时间分辨薄膜形成实验组织为可解释的潜结构。与DINOv3(ViT-7B)等通用视觉基础模型对比表明,领域专用训练能带来更可解释的潜空间组织。两种工作流已集成至MLExchange平台中的潜空间探索器(Latent Space Explorer),支持对归档数据集和实时实验的交互式结构探索。
原文摘要 · Abstract (English)
Scientific user facilities generate X-ray scattering data faster than traditional workflows can process them. We address this challenge across two settings, offline dataset exploration and live on-the-fly analysis. We train a domain-specific attention-based Convolutional Variational Autoencoder (C-VAE) on 1.5 million X-ray scattering images to learn low-dimensional representations capturing structural variation across diverse experimental conditions. The learned latent space reveals well-organized clusters and smooth trajectories reflecting experimental progression. It further supports controlled synthetic scattering image generation across diverse structural states. When deployed without retraining, the model organizes time-resolved film formation experiments at two synchrotron facilities into interpretable latent structures. Benchmarking against DINOv3 (ViT-7B), a general-purpose vision foundation model, demonstrates that domain-specific training yields more interpretable latent organization for scattering data. Both workflows are integrated within Latent Space Explorer, a component of the MLExchange platform, supporting interactive structural exploration across archived datasets and live experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。