arXiv:2602.01710cs.CVcond-mat.mtrl-sci2026-02

用物理仿真+生成模型,零标注实现显微图像精准分割。

Physics Informed Generative AI Enabling Labour Free Segmentation For Microscopy Analysis

  • 用相场模拟生成带真实标签的微结构数据,再通过CycleGAN转为逼真扫描电镜图像。
  • 模型在真实实验图像上达0.90边界F1分数和0.88交并比,表现接近人工标注水平。
  • 适合材料科学领域研究者,可大幅降低数据标注成本,加速新材料发现。

显微图像语义分割对高通量材料表征至关重要,但其自动化受限于专家标注数据的高昂成本、主观性及稀缺性。尽管基于物理的仿真可提供可扩展的替代方案,但以往模型因存在显著域差距,难以泛化——缺乏真实数据中的复杂纹理、噪声模式和成像伪影。本文提出一种无需人工标注的分割框架,成功弥合仿真到现实的差距。该流程利用相场模拟生成大量具有完美内生真值掩码的微结构形态;随后采用无配对图像到图像翻译的循环一致生成对抗网络(CycleGAN),将清晰仿真图像转换为大规模高保真真实扫描电镜图像。仅在该合成数据上训练的U-Net模型,在未见的真实实验图像上展现出优异泛化能力,平均边界F1得分为0.90,交并比(IOU)达0.88。通过t-SNE特征空间投影与香农熵分析,证实合成图像在统计和特征层面与真实数据分布不可区分。本框架彻底摆脱手动标注依赖,将数据稀缺问题转化为数据丰沛问题,为材料发现与分析提供稳健、全自动的解决方案。

原文摘要 · Abstract (English)

Semantic segmentation of microscopy images is a critical task for high-throughput materials characterisation, yet its automation is severely constrained by the prohibitive cost, subjectivity, and scarcity of expert-annotated data. While physics-based simulations offer a scalable alternative to manual labelling, models trained on such data historically fail to generalise due to a significant domain gap, lacking the complex textures, noise patterns, and imaging artefacts inherent to experimental data. This paper introduces a novel framework for labour-free segmentation that successfully bridges this simulation-to-reality gap. Our pipeline leverages phase-field simulations to generate an abundant source of microstructural morphologies with perfect, intrinsically-derived ground-truth masks. We then employ a Cycle-Consistent Generative Adversarial Network (CycleGAN) for unpaired image-to-image translation, transforming the clean simulations into a large-scale dataset of high-fidelity, realistic SEM images. A U-Net model, trained exclusively on this synthetic data, demonstrated remarkable generalisation when deployed on unseen experimental images, achieving a mean Boundary F1-Score of 0.90 and an Intersection over Union (IOU) of 0.88. Comprehensive validation using t-SNE feature-space projection and Shannon entropy analysis confirms that our synthetic images are statistically and featurally indistinguishable from the real data manifold. By completely decoupling model training from manual annotation, our generative framework transforms a data-scarce problem into one of data abundance, providing a robust and fully automated solution to accelerate materials discovery and analysis.

显微分割生成模型材料科学零标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。