机器学习算子无法零样本实现超分辨率,需新训练方法
The False Promise of Zero-Shot Super-Resolution in Machine-Learned Operators
- 将多分辨率推理拆解为频率外推与分辨率插值两部分
- 实验证明模型在未训练的高分辨率上会严重失真并产生伪影
- 提出高效数据驱动训练法,解决伪影问题并提升泛化能力
科学机器学习的核心挑战在于建模连续现象,而这些现象实际以离散形式表示。机器学习算子(MLOs)被提出用于实现这一目标,因其可在任意分辨率下进行推理。本文评估了该架构是否足以实现“零样本超分辨率”——即在未训练过的更高分辨率数据上进行准确推理。我们全面检验了MLOs在零样本下的亚分辨率与超分辨率推理表现。将多分辨率推理解耦为两个关键行为:1)对不同频率信息的外推;2)在不同分辨率间的插值。实证表明,MLOs在零样本情况下均无法有效完成这两项任务。因此,MLOs在不同于训练分辨率的条件下无法实现准确推理,反而表现出脆弱性且易受混叠影响。为此,我们提出一种简单、计算高效且数据驱动的多分辨率训练协议,可克服混叠问题,并实现稳健的多分辨率泛化。
原文摘要 · Abstract (English)
A core challenge in scientific machine learning, and scientific computing more generally, is modeling continuous phenomena which (in practice) are represented discretely. Machine-learned operators (MLOs) have been introduced as a means to achieve this modeling goal, as this class of architecture can perform inference at arbitrary resolution. In this work, we evaluate whether this architectural innovation is sufficient to perform "zero-shot super-resolution," namely to enable a model to serve inference on higher-resolution data than that on which it was originally trained. We comprehensively evaluate both zero-shot sub-resolution and super-resolution (i.e., multi-resolution) inference in MLOs. We decouple multi-resolution inference into two key behaviors: 1) extrapolation to varying frequency information; and 2) interpolating across varying resolutions. We empirically demonstrate that MLOs fail to do both of these tasks in a zero-shot manner. Consequently, we find MLOs are not able to perform accurate inference at resolutions different from those on which they were trained, and instead they are brittle and susceptible to aliasing. To address these failure modes, we propose a simple, computationally-efficient, and data-driven multi-resolution training protocol that overcomes aliasing and that provides robust multi-resolution generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。