提出无需训练的机器学习新方法,解决函数逼近与迁移学习中的理论难题。
Learning Without Training
- 基于数学理论改进监督学习的函数逼近机制
- 揭示函数在部分数据下的可延拓性与平滑性关系
- 融合信号分离思想提升分类效率,适合快速主动学习场景
机器学习是应对海量数据实际问题的核心。随着神经网络在大规模问题上的成功,机器学习研究日益活跃。本文聚焦三个基于数学理论的机器学习应用项目。第一项关注监督学习与流形学习,针对函数逼近问题(给定数据集 $\\(mathcal{D}=\{(x_j,f(x_j))\}_{j=1}^M$,构建近似模型 $F\approx f$),提出一种可弥补当前范式理论缺陷的新方法。第二项研究迁移学习,探讨在仅知部分域数据时,如何将源域的近似过程或模型用于目标域的优化,重点分析函数延拓在目标空间子集上的定义条件,以及局部光滑性与延拓后光滑性的关联。第三项聚焦分类任务,尤其在主动学习框架下,提出一种源自信号分离技术的替代方法,建立信号分离与分类的统一理论,并设计出在精度上媲美最新算法、但运行速度显著更快的新算法。
原文摘要 · Abstract (English)
Machine learning is at the heart of managing the real-world problems associated with massive data. With the success of neural networks on such large-scale problems, more research in machine learning is being conducted now than ever before. This dissertation focuses on three different projects rooted in mathematical theory for machine learning applications. The first project deals with supervised learning and manifold learning. In theory, one of the main problems in supervised learning is that of function approximation: that is, given some data set $\mathcal{D}=\{(x_j,f(x_j))\}_{j=1}^M$, can one build a model $F\approx f$? We introduce a method which aims to remedy several of the theoretical shortcomings of the current paradigm for supervised learning. The second project deals with transfer learning, which is the study of how an approximation process or model learned on one domain can be leveraged to improve the approximation on another domain. We study such liftings of functions when the data is assumed to be known only on a part of the whole domain. We are interested in determining subsets of the target data space on which the lifting can be defined, and how the local smoothness of the function and its lifting are related. The third project is concerned with the classification task in machine learning, particularly in the active learning paradigm. Classification has often been treated as an approximation problem as well, but we propose an alternative approach leveraging techniques originally introduced for signal separation problems. We introduce theory to unify signal separation with classification and a new algorithm which yields competitive accuracy to other recent active learning algorithms while providing results much faster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。