揭示神经算子零样本超分辨率的理论边界,发现其并非总可行
Is Zero-Shot Super-Resolution Possible in Operator Learning?
- 从信息论角度证明:即使输入完整,零样本超分也可能不可能
- 提出输出函数霍尔德光滑性为超分辨率的充分条件
- 实验验证理论预测的失败模式,适合研究算子学习理论者
神经算子常被报道具备零样本超分辨率能力,即在粗网格上训练的模型无需再训练即可准确预测细网格结果。尽管有强烈实证证据,但该现象的理论基础仍不明确。本文系统研究了算子学习中的零样本超分辨率。首先,我们证明即使在理想场景下(输入函数在整个连续域上已知,真值为简单秩一线性算子),零样本超分辨率在信息论上仍可能不可行。随后,我们识别出输出函数的霍尔德光滑性是实现零样本超分辨率的充分条件,并推导出相应的泛化界。最后,通过实验验证了所提出的失效模式。
原文摘要 · Abstract (English)
Neural operators are often reported to exhibit zero-shot super-resolution, a phenomenon in which a model trained on coarse grids produces accurate predictions on finer testing grids without additional retraining. Despite strong empirical evidence, the theoretical foundations of this phenomenon remain unclear. In this work, we provide a systematic theoretical study of zero-shot super-resolution in operator learning. We first show that zero-shot super-resolution can be information-theoretically impossible even in benign settings such as when the input functions are available over the entire continuum and the ground truth is a simple rank-one linear operator. We then identify H{\" o}lder smoothness of the output functions as a sufficient condition for zero-shot super-resolution and derive corresponding generalization bounds. Finally, we also validate the identified failure modes through experimental results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。