arXiv:2504.06991cs.LGmath.PR2025-04中稿 · publication in San…被引 2

研究数据分批时如何控制批次内相似度,提升学习效率。

Dissimilar Batch Decompositions of Random Datasets

  • 从概率角度设计分批策略,限制每批内数据相似度。
  • 推导出在高概率下最小批次大小的理论边界。
  • 用鞅方法分析相似数据子集的最大可能规模,适合理论研究者。

为提升学习效果,大型数据集常被拆分为小批次依次输入预测模型。本文从概率视角研究此类分批方式:假设数据点(可能含噪声)独立来自给定空间,并定义两点间的相似度概念。研究受限于每批内相似度的分批方案,推导出最小批次大小的高概率上界。结果揭示了放宽相似度约束与整体批次规模之间的内在权衡;同时利用鞅方法,获得具有特定相似度的数据子集最大规模的上界。

原文摘要 · Abstract (English)

For better learning, large datasets are often split into small batches and fed sequentially to the predictive model. In this paper, we study such batch decompositions from a probabilistic perspective. We assume that data points (possibly corrupted) are drawn independently from a given space and define a concept of similarity between two data points. We then consider decompositions that restrict the amount of similarity within each batch and obtain high probability bounds for the minimum size. We demonstrate an inherent tradeoff between relaxing the similarity constraint and the overall size and also use martingale methods to obtain bounds for the maximum size of data subsets with a given similarity.

数据分批概率分析理论边界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。