代码生成比强化学习更易突破,因反馈信息更密集可验证。
Why Code, Why Now: An Information-Theoretic Perspective on the Limits of Machine Learning
- 用信息结构分级评估任务可学习性,超越模型大小与算法选择。
- 提出五级可学习性层级,解释为何代码任务能稳定扩展。
- 适合研究机器学习边界、系统设计的学者参考。
本文从信息论视角重新审视机器学习的局限性:进展上限不取决于模型规模或算法选择,而由任务本身的信息结构决定。代码生成比强化学习进展更可靠,主要因为代码在每个词元层面提供密集、局部且可验证的反馈,而多数强化学习任务缺乏此类高质量反馈。这种反馈质量差异并非二元,而是连续分级的。本文提出基于信息结构的五级可学习性层级,并论证诊断任务在此层级中的位置,比任何模型属性更能预测规模化效果。该框架建立在计算问题三个属性(表达性、可计算性、可学习性)的严格区分之上,揭示其两两关系及蕴含关系的适用边界,提供统一模板以明确结构性差异。分析表明,监督学习在代码任务上可预测地扩展,而强化学习则不然;同时,单纯依赖规模解决剩余挑战的假设值得重新审视。
原文摘要 · Abstract (English)
This paper offers a new perspective on the limits of machine learning: the ceiling on progress is set not by model size or algorithm choice but by the information structure of the task itself. Code generation has progressed more reliably than reinforcement learning, largely because code provides dense, local, verifiable feedback at every token, whereas most reinforcement learning problems do not. This difference in feedback quality is not binary but graded. We propose a five-level hierarchy of learnability based on information structure and argue that diagnosing a task's position in this hierarchy is more predictive of scaling outcomes than any property of the model. The hierarchy rests on a formal distinction among three properties of computational problems (expressibility, computability, and learnability). We establish their pairwise relationships, including where implications hold and where they fail, and present a unified template that makes the structural differences explicit. The analysis suggests why supervised learning on code scales predictably while reinforcement learning does not, and why the common assumption that scaling alone will solve remaining ML challenges warrants scrutiny.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。