arXiv:2604.21903cs.LGcs.AI2026-04

提出可跨尺度的时空超分辨率框架,统一处理不同分辨率与帧率需求。

A Scale-Adaptive Framework for Joint Spatiotemporal Super-Resolution with Diffusion Models

论文配图:A Scale-Adaptive Framework for Joint Spatiotemporal Super-Resolution with Diffusion Models
图 1 · 摘自论文原文
  • 将时空超分辨率分解为确定性预测与残差扩散模型,实现跨尺度复用。
  • 在法国再分析降水数据上,支持空间1-25倍、时间1-6倍的联合超分辨率。
  • 通过调节噪声、上下文长度等参数,适配不同尺度,适合气候建模应用。

深度学习视频超分辨率发展迅速,但气候应用通常仅提升空间或时间分辨率,联合时空模型常针对单一放大因子设计,难以跨分辨率和帧率迁移。本文提出一种尺度自适应框架,通过分解时空超分辨率为确定性条件均值预测(带注意力)与残差条件扩散模型,并引入可选的质量守恒变换以保持总降水量一致。假设大放大因子主要增加不确定性而非改变均值结构,通过调整三个因子相关超参数即可实现尺度自适应:扩散噪声调度幅度β(放大因子越大,β越大以增强多样性)、时间上下文长度L(保持不同帧率下的注意力视野一致),以及可选的质量守恒函数f(对大因子进行衰减,抑制极端值放大)。在法国再分析降水数据集Comephore上验证,同一架构可覆盖空间放大1-25倍、时间放大1-6倍的范围,实现跨尺度通用的联合时空超分辨率。

原文摘要 · Abstract (English)

Deep-learning video super-resolution has progressed rapidly, but climate applications typically super-resolve (increase resolution) either space or time, and joint spatiotemporal models are often designed for a single pair of super-resolution (SR) factors (upscaling spatial and temporal ratio between the low-resolution sequence and the high-resolution sequence), limiting transfer across spatial resolutions and temporal cadences (frame rates). We present a scale-adaptive framework that reuses the same architecture across factors by decomposing spatiotemporal SR into a deterministic prediction of the conditional mean, with attention, and a residual conditional diffusion model, with an optional mass-conservation (same precipitation amount in inputs and outputs) transform to preserve aggregated totals. Assuming that larger SR factors primarily increase underdetermination (hence required context and residual uncertainty) rather than changing the conditional-mean structure, scale adaptivity is achieved by retuning three factor-dependent hyperparameters before retraining: the diffusion noise schedule amplitude beta (larger for larger factors to increase diversity), the temporal context length L (set to maintain comparable attention horizons across cadences) and optionally a third, the mass-conservation function f (tapered to limit the amplification of extremes for large factors). Demonstrated on reanalysis precipitation over France (Comephore), the same architecture spans super-resolution factors from 1 to 25 in space and 1 to 6 in time, yielding a reusable architecture and tuning recipe for joint spatiotemporal super-resolution across scales.

时空超分扩散模型气候建模尺度自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。