arXiv:2507.06738cs.CVcs.AI2025-07

提出DIFFUMA模型与晶圆切割数据集,提升工业视频预测精度

DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement

  • 双路径架构:Mamba捕获长时序全局信息,扩散模块增强细节
  • 在自建数据集上MSE降低39%,SSIM提升至0.988,接近完美
  • 首次公开半导体晶圆切割视频数据集,助力工业AI研究

时空视频预测在气象预报、工业自动化等领域至关重要。然而,在半导体制造等高精度场景中,缺乏专用基准数据集严重制约了复杂过程建模与预测的研究。为此,我们做出两方面贡献:首先,构建并发布首个面向半导体晶圆切割工艺的公开时间图像数据集——芯片切割流水线数据集(CHDL),该数据集由工业级视觉系统采集,为高保真过程建模、缺陷检测和数字孪生开发提供了亟需且具有挑战性的基准。其次,提出DIFFUMA,一种专为细粒度动态设计的双路径预测架构。模型通过并行的Mamba模块捕捉全局长时序上下文,同时利用受时序特征引导的扩散模块恢复并增强细粒度空间细节,有效缓解特征退化问题。实验表明,在所提出的CHDL基准上,DIFFUMA显著优于现有方法,将均方误差(MSE)降低39%,结构相似性(SSIM)从0.926提升至近乎完美的0.988。该优异性能在自然现象数据集上也具泛化能力。本工作不仅带来新的最先进(SOTA)模型,更向社区提供宝贵的数据资源,推动工业AI未来发展。

原文摘要 · Abstract (English)

Spatio-temporal video prediction plays a pivotal role in critical domains, ranging from weather forecasting to industrial automation. However, in high-precision industrial scenarios such as semiconductor manufacturing, the absence of specialized benchmark datasets severely hampers research on modeling and predicting complex processes. To address this challenge, we make a twofold contribution.First, we construct and release the Chip Dicing Lane Dataset (CHDL), the first public temporal image dataset dedicated to the semiconductor wafer dicing process. Captured via an industrial-grade vision system, CHDL provides a much-needed and challenging benchmark for high-fidelity process modeling, defect detection, and digital twin development.Second, we propose DIFFUMA, an innovative dual-path prediction architecture specifically designed for such fine-grained dynamics. The model captures global long-range temporal context through a parallel Mamba module, while simultaneously leveraging a diffusion module, guided by temporal features, to restore and enhance fine-grained spatial details, effectively combating feature degradation. Experiments demonstrate that on our CHDL benchmark, DIFFUMA significantly outperforms existing methods, reducing the Mean Squared Error (MSE) by 39% and improving the Structural Similarity (SSIM) from 0.926 to a near-perfect 0.988. This superior performance also generalizes to natural phenomena datasets. Our work not only delivers a new state-of-the-art (SOTA) model but, more importantly, provides the community with an invaluable data resource to drive future research in industrial AI.

视频预测工业AI扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。