提出跨维度迁移的理论框架,解决模型在不同输入尺寸间的泛化问题。
On Transferring Transferability: Towards a Theory for Size Generalization
- 基于极限空间中的连续性定义跨维度迁移能力
- 验证了现有架构在调整后可实现尺寸无关性能
- 为设计新模型提供可迁移的设计原则
许多现代学习任务需要处理不同尺寸的输入。为此,针对图、集合和点云等领域的维度无关架构被提出。近期关于图神经网络的研究探讨了在低维数据上训练的模型能否迁移到高维输入。本文通过引入一个通用的跨维度迁移框架,扩展了该研究方向。我们证明,迁移能力恰好对应于将小规模问题实例与等价大规模实例识别后的极限空间中的连续性。这种识别由数据和学习任务驱动。我们在现有架构上实例化该框架,并实施必要修改以确保其迁移能力。最后,我们提出了设计新可迁移模型的原则。数值实验验证了我们的发现。
原文摘要 · Abstract (English)
Many modern learning tasks require models that can take inputs of varying sizes. Consequently, dimension-independent architectures have been proposed for domains where the inputs are graphs, sets, and point clouds. Recent work on graph neural networks has explored whether a model trained on low-dimensional data can transfer its performance to higher-dimensional inputs. We extend this body of work by introducing a general framework for transferability across dimensions. We show that transferability corresponds precisely to continuity in a limit space formed by identifying small problem instances with equivalent large ones. This identification is driven by the data and the learning task. We instantiate our framework on existing architectures, and implement the necessary changes to ensure their transferability. Finally, we provide design principles for designing new transferable models. Numerical experiments support our findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。