用几何熵混合优化大模型数据配比,提升训练效果。
GEM: Geometric Entropy Mixing for Optimal LLM Data Curation

- 在超球面上建模数据混合,用变分方法求解最优配比。
- 实验显示下游任务平均准确率提升1.2%,优于现有策略。
- 适合关注数据质量与可解释性配比的研究者使用。
大模型预训练效果越来越依赖数据组成而非单纯数量。然而,最优混合因分类缺陷受阻:人类分类存在本体论错位,欧氏聚类无法解决嵌入的各向异性。本文提出GEM(几何熵混合),将数据筛选重构为在超球面的变分问题,并引入混合平衡正则项。通过分离生成先验并使用可证明收敛的MM(极小化-极大化)算法,有效防止聚类坍塌,发现欧氏启发法无法察觉的平衡语义结构。采用教师-学生蒸馏技术将此几何保真度扩展至网络规模语料库,并提出几何影响得分(GIS)实现可解释的分类体系生成。在11亿参数模型上的实验表明,将GEM整合进DoReMi和RegMix等混合策略中,可建立新基准,平均下游准确率提升达1.2%,并提供可预测的数据混合坐标系统。
原文摘要 · Abstract (English)
LLM pre-training efficacy increasingly depends on data composition rather than sheer volume. Yet, optimal mixing is hindered by categorization flaws: human taxonomies suffer from ontological misalignment, and Euclidean clustering fails to address embedding anisotropy. We introduce GEM (Geometric Entropy Mixing), a framework reformulating data curation as a variational problem on the hypersphere augmented with a mixing-balance regularizer. By decoupling the generative prior and optimizing the objective via a provable MM (Minorize-Maximize) algorithm, GEM effectively counteracts the cluster collapse to discover balanced semantic structures invisible to Euclidean heuristics. We employ teacher-student distillation to scale this geometric fidelity to web-scale corpora and introduce the Geometric Influence Score (GIS) for interpretable taxonomy generation. Experiments with 1.1B-parameter models demonstrate that GEM establishes a new state-of-the-art when integrated into mixing strategies like DoReMi and RegMix, improving average downstream accuracy by up to 1.2% and offering a robust coordinate system for predictable data mixing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。