改进统计学习理论中测度性条件,使学习可证性更宽松
Null Measurability at the Symmetrization Interface in VC Learning
- 在对称化证明的单侧鬼差界面,弱化测度性要求为解析集
- 构造反例证明弱条件严格优于传统Borel条件
- 新条件在概念类拼接、插值等操作下保持稳定
近期工作重新审视统计学习基本定理中的可测性问题,要求鬼差上确界具有Borel可测性。本文指出,在标准对称化证明所用的单侧鬼差界面,该要求过强。对于任意定义在波兰空间上的Borel参数化概念类,事件'存在一个假设其鬼经验误差比训练经验误差至少大ε/2'是解析集。由Choquet容度定理,该事件在每个有限Borel测度的完备化下均可测。进一步构造出一个坏事件为零测但非Borel的概念类,实现与Borel上确界条件的严格分离。最后证明该弱正则性在固定和可数插值、纤维积合并等自然概念类构造下保持封闭。在可实现情形下,这些结果弱化了从有限VC维到PAC可学习性所需的对称化路径中的可测性假设。主要结果及所用描述集合论工具已在Lean 4中形式化。
原文摘要 · Abstract (English)
Recent work revisiting measurability in the fundamental theorem of statistical learning imposes Borel measurability of ghost-gap suprema. We show that, at the one-sided ghost-gap interface actually used by the standard symmetrization proof, this requirement is stronger than necessary. For any Borel-parameterized concept class on a Polish domain, the bad event "there exists a hypothesis whose ghost empirical error exceeds its training empirical error by at least ε/2" is analytic. By Choquet capacitability, it is therefore measurable in the completion of every finite Borel measure. We then construct a concept class whose bad event is null-measurable but not Borel, giving a strict separation from the Borel supremum condition. Finally, we prove closure under patching, fixed and countable interpolation, and fiber-product amalgamation, showing that the weaker regularity level is stable under natural concept-class constructors. In the realizable setting, where targets belong to the class and are measurable, these results weaken the measurability hypothesis needed by the symmetrization route from finite VC dimension to PAC learnability. The main results and the descriptive-set-theoretic infrastructure used by them are formalized in Lean 4.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。