别盲目堆数据,要选对任务再扩数据。
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling
- 根据数据的组合与结构特征决定扩数据优先级
- 某些任务即使多数据也难提升性能
- 适合关注模型效率与数据质量的研究者
尽管大语言模型需要越来越多的数据来训练和扩展,但我们不应随意获取任何数据,而应有意识地选择更可能从数据扩展中受益的任务。我们主张在数据获取上保持审慎,因为数据本身的结构与组合模式决定了哪些任务值得优先扩展数据。这些模式还影响下一代计算范式的发展,尤其针对数据扩展效率低或无效的任务。
原文摘要 · Abstract (English)
While Large Language Models require more and more data to train and scale, rather than looking for any data to acquire, we should consider what types of tasks are more likely to benefit from data scaling. We should be intentional in our data acquisition. We argue that the shape of the data itself, such as its compositional and structural patterns, informs which tasks to prioritize in data scaling, and shapes the development of the next generation of compute paradigms for tasks where data scaling is inefficient, or even insufficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。