arXiv:2607.01010stat.MLcs.IT2026-07被引 2

为低维数据分类提供数学框架,揭示数据结构对模型能力的影响。

Function-Counting Theory for Low-Dimensional Data Structures

  • 基于柯弗理论改进假设,考虑数据低维特性
  • 推导反映数据结构的分类二分计数方法
  • 适用于研究数据结构如何影响模型性能

深度学习在分类与回归任务中的成功,常归因于真实世界数据虽高维表征却具有低维结构。本文旨在为低维数据上的二分类问题建立数学框架,基于柯弗(1965)的函数计数理论。原理论依赖一般位置假设,忽略数据内在结构。本文修正该假设以体现数据的低维性,推导出反映数据结构的分类二分计数。进一步将柯弗的分离能力与泛化问题拓展至低维设定,使数据结构对这两方面的影响得以分析。

原文摘要 · Abstract (English)

The success of deep learning models in classification and regression is widely attributed to the low-dimensional structure that real-world data tend to exhibit, despite their high-dimensional representation. This work attempts to provide a mathematical framework for binary classification on low-dimensional data, building on Cover's (1965) function-counting theory. With our framework, we aim to address the question of how the low-dimensional structure of the data affects the classification capabilities of learning models. Cover's theory relies on a general position assumption that blinds it to the underlying data structure. We refine this assumption to account for the low-dimensionality of the data and derive dichotomy counts that reflect the data structure. We further extend Cover's separation capacity and problem of generalization to the low-dimensional setting, enabling the impact of the underlying data structure on both to be analyzed.

分类理论低维数据函数计数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。