arXiv:2510.02456cs.LGcs.AI2025-10被引 1

用市场机制统一整合多种数据价值信号,高效选出高性价比训练样本。

Market-Driven Subset Selection for Budgeted Training

  • 将每个样本视为可交易合约,用市场规则融合不确定性、稀有性等多维价值信号。
  • 在60k token预算下,数学推理任务性能媲美强基线,方差更低且开销小于0.1 GPU小时。
  • 适合资源受限场景,尤其对提示级推理与分类任务的高效数据筛选有启发。

在大规模数据上训练大语言模型计算成本高昂,但实证表明大量样本对最终性能贡献甚微。数据子集选择通过在资源约束下识别小而高价值的子集来提升效率。然而,样本价值具有多重属性,如不确定性、分布稀有性和多样性,这些异质信号通常通过无理论依据的加权求和组合。我们提出一种基于市场的框架,将每个训练样本视为可交易合约,利用对数市场评分规则将多维价值信号聚合为统一价格。异质信号作为“交易者”,单一流动性参数控制集中度与平滑性,主题级归一化确保校准聚合。通过每令牌价格决策规则显式处理令牌预算,并引入可解释的长度偏差参数。我们建立了与最大熵聚合的理论关联,并在噪声但单调信号下提供价值恢复保证。在严格60k令牌预算下的GSM8K数学推理任务中,本方法性能媲美强单信号基线,方差更低,开销低于0.1 GPU小时;在AGNews分类任务中,5%-25%保留率下达到竞争性准确率并提升稳定性。该框架在固定计算预算下统一了多信号数据筛选,适用于提示级推理与分类任务。

原文摘要 · Abstract (English)

Training large language models on massive datasets is computationally expensive, yet empirical evidence suggests that substantial portions of training examples contribute minimally to final performance. Data subset selection addresses this inefficiency by identifying small, high-utility subsets under resource constraints. However, example utility is inherently multi-faceted, encompassing uncertainty, distributional rarity, and diversity signals that are heterogeneous and typically combined through ad hoc weighted sums lacking theoretical grounding. We propose a market-based framework that treats each training example as a tradeable contract and employs the Logarithmic Market Scoring Rule to aggregate multiple utility signals into coherent prices. Heterogeneous signals act as traders, a single liquidity parameter controls concentration versus smoothing, and topic-wise normalization ensures calibrated aggregation. Token budgets are handled explicitly through a price-per-token decision rule with an interpretable length-bias parameter. We establish theoretical connections to maximum-entropy aggregation and provide utility recovery guarantees under noisy but monotone signals. On GSM8K mathematical reasoning under strict 60k-token budgets, our selector achieves parity with strong single-signal baselines while exhibiting lower variance and incurring less than 0.1 GPU-hour overhead. On AGNews classification at 5-25\% retention rates, the market formulation delivers competitive accuracy with improved stability. Our framework unifies multi-signal data curation under fixed computational budgets for prompt-level reasoning and classification tasks.

数据筛选市场机制预算约束多信号融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。