arXiv:2604.07363cs.LG2026-04

数据分布影响大模型学习,好成绩未必代表真能力。

Benchmark Shadows: Data Alignment, Parameter Footprints, and Generalization in Large Language Models

  • 用可控数据干预分离分布影响,对比不同数据效果。
  • 对齐基准的数据提升局部指标,但限制泛化能力。
  • 多模型验证发现此现象普遍,适合关注模型真实能力的研究者。

大型语言模型在基准测试中表现优异,却未必具备更广泛的能力。我们假设这种差异源于训练过程中数据分布带来的训练制度差异。为此,我们设计了受控的数据干预,在固定训练设置下隔离分布效应。结果表明,与基准对齐的数据能提升特定评估指标,但抑制了更广泛的表征发展;而扩大覆盖范围的数据则引发更分散的参数适应,带来更好的泛化性能。我们进一步引入基于谱分析和秩分析的参数空间诊断方法,揭示了这些训练模式的独特结构特征。该现象在多种开源模型族(包括多模态模型)中均被观察到,说明其不局限于受控环境。针对提示重复的案例研究显示,并非所有数据缺陷都会引发训练模式转变。这些结果表明,仅凭基准表现无法全面刻画模型能力,强调了数据分布对学习动态的关键作用。

原文摘要 · Abstract (English)

Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in training regimes induced by data distribution. To investigate this, we design controlled data interventions that isolate distributional effects under fixed training settings. We find that benchmark-aligned data improves narrow evaluation metrics while limiting broader representational development, whereas coverage-expanding data leads to more distributed parameter adaptation and better generalization. We further introduce parameter-space diagnostics based on spectral and rank analyses, which reveal distinct structural signatures of these regimes. Similar patterns are observed across diverse open-source model families, including multimodal models as a key case study, suggesting that these effects extend beyond controlled settings. A case study on prompt repetition shows that not all data artifacts induce regime shifts. These results indicate that benchmark performance alone is insufficient to characterize model capability, and highlight the importance of data distribution in shaping learning dynamics.

大模型数据分布泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。