arXiv:2511.21692cs.CLcs.AI2025-11被引 6

大模型在难易任务间泛化能力有限,难易数据都不可靠。

Revisiting Generalization Across Difficulty Levels: It's Not So Easy

  • 用多模型和项目反应理论客观评估六大数据集难度
  • 无论训练用简单或复杂数据,泛化效果均不一致
  • 适合需要全面覆盖难度的评测与训练场景

我们系统评估了大语言模型在不同任务难度间的泛化能力,这是数据筛选与评估的关键问题。现有研究对训练于简单或复杂数据是否更优、增益能否在难易测试上体现存在分歧。本研究通过数千个不同大模型的输出与项目反应理论(IRT),对六个数据集中的样本进行难度排序,避免人为判断影响。结果表明,跨难度泛化常受限:无论训练数据是简单还是复杂,都无法在全难度范围内实现稳定提升。这说明训练与评估数据都需包含多样难度,简化难度处理风险极高。

原文摘要 · Abstract (English)

We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation. Existing research is mixed regarding whether training on easier or harder data leads to better results, and whether those gains come on easier or harder test data. We address this question by conducting a systematic evaluation of LLMs' generalization across models, datasets, and fine-grained groups of example difficulty. We rank examples in six datasets using the outputs of thousands of different LLMs and Item Response Theory (IRT), a well-established difficulty metric in educational testing. Unlike prior work, our difficulty ratings are therefore determined solely by the abilities of many different LLMs, excluding human opinions of difficulty. With a more objective, larger-scale, and finer-grained analysis, we show that cross-difficulty generalization is often limited; training on either easy or hard data cannot achieve consistent improvements across the full range of difficulties. These results show the importance of having a range of difficulties in both training and evaluation data for LLMs, and that taking shortcuts with respect to difficulty is risky.

大模型泛化能力难度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。