arXiv:2608.04030cs.GRcs.AI2026-08

用专业数据微调扩散模型,让AI更准确生成核能概念图

NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts

论文配图:NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts
图 1 · 摘自论文原文
  • 用1000张带标签的核能图像微调开源扩散模型
  • 微调后模型在工程细节上准确性显著提升,尤其对复杂概念
  • 适合核能教育、科研人员使用,提升AI生成的专业可信度

生成式人工智能已革新文本到图像合成,但在专业工程领域仍缺乏探索。以核工程为例,通用基础模型常生成物理错误或概念不一致的图像,因缺乏领域知识。本文首次系统研究通过微调开源扩散模型实现核能领域的文本到图像生成。我们构建了一个包含1000张带标注图像的数据集,涵盖反应堆、燃料循环、辐射等概念,并用于微调三种先进开源模型:Stable Diffusion XL (SDXL)、SD-v3.5-Medium 和 flow-matching Flux.1。通过定量图像相似性指标与专家定性评估,对比微调前后及零样本模型表现。结果表明,微调显著提升了SDXL的保真度,对SD-v3.5-Medium效果有限,对Flux.1无明显改善,说明适应效果高度依赖生成架构而非仅模型规模。进一步对比GPT-Image-2、Gemini-3.1-Flash-Image和Midjourney三款主流商用系统,虽后者在通用核能概念上表现良好,但在专业工程提示下频繁出错,而微调后的开源模型生成更准确、技术一致的结果。研究证实,领域特异性微调是构建可信生成式AI工具的有效路径。

原文摘要 · Abstract (English)

Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains largely unexplored. As an exmaple in nuclear engineering, general-purpose foundation models frequently generate physically incorrect or conceptually inconsistent images because they lack domain-specific knowledge. This work presents one of the first systematic studies of domain adaptation for nuclear text-to-image generation through fine-tuning of open-source diffusion models. We curate a dataset of 1,000 captioned nuclear energy images spanning reactors, fuel cycles, radiation, and related concepts, and use it to fine-tune three state-of-the-art open-source models: Stable Diffusion XL (SDXL), SD-v3.5-Medium, and the flow-matching Flux.1 model. Their performance is evaluated using both quantitative image-similarity metrics and qualitative expert assessment against the corresponding zero-shot models. Fine-tuning substantially improves the fidelity of SDXL, provides only limited gains for SD-v3.5-Medium, and yields no measurable improvement for Flux.1, demonstrating that adaptation effectiveness depends strongly on the underlying generative architecture rather than model scale alone. We further compare the fine-tuned models against three leading commercial systems--GPT-Image-2, Gemini-3.1-Flash-Image, and Midjourney. Although GPT-Image-2 and Gemini generate convincing images for broad nuclear concepts, they frequently fail on specialized engineering prompts, where the fine-tuned open-source models produce more accurate and technically consistent outputs. These results establish domain-specific fine-tuning as a practical pathway for developing trustworthy generative AI tools for domain-specific applications.

文本生成图像核能领域微调扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。