arXiv:2607.07680math.STcs.LG2026-07

用随机采样统一解决不同尺寸输入的泛化与压缩问题

Any-Dimensional Learning by Sampling

  • 设计随机采样映射,实现跨尺寸输入的可比性与近似
  • 给出函数类在采样连续性下的泛化与压缩速率
  • 适用于序列、图、张量等多类模型,尤其适合对称结构

许多机器学习模型需处理不同规模的输入,如点云、变长序列和异构图。这些模型仅在有限数量且规模受限的样本上训练。如何使模型从较小输入推广到训练中未见的大规模输入?同时,大输入评估成本高昂,能否通过采样得到小规模近似输入以保留模型输出?核心在于比较不同规模输入并近似大输入。本文提出统一框架,利用随机采样映射(如重采样、随机分箱、物种采样)解决上述问题。根据不同问题场景中实例间的对称性与关系,界定各类采样的适用范围。该框架为在特定采样意义下连续的函数类提供明确的泛化与压缩速率,涵盖序列、图与张量上的大量函数,包括测度上的矩多项式、同态密度、图计数、置换不变变换器及图神经网络。

原文摘要 · Abstract (English)

Many machine learning models are defined for inputs of different sizes, such as point clouds containing different numbers of points, sequences of tokens of different lengths, and graphs on different numbers of nodes. Such models are trained on finitely-many examples of necessarily limited sizes. How well do these models generalize from inputs of small size to larger inputs of size not seen during training? Furthermore, evaluating such models on large inputs is often expensive. How can we sketch large inputs to obtain smaller ones on which the model takes similar values? At the heart of both questions is the need to compare inputs of different sizes and to approximate large inputs by small ones. We present a unified approach to address these questions by using random sampling maps to compare inputs of different sizes. The sampling maps we consider are generalizations of sampling with replacement, random binning, and species sampling. We characterize the application domains in which each type of sampling is appropriate in terms of the symmetries and relations between problem instances of different sizes in the domain. Our framework yields explicit generalization and sketching rates for function classes continuous with respect to a chosen notion of sampling, encompassing large families of functions defined on sequences, graphs, and tensors of different sizes. Specific examples include moment polynomials on measures, homomorphism densities and numbers of graphs, permutation-invariant transformers, and graph neural networks.

泛化分析采样方法模型压缩图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。