arXiv:2505.04075cs.LGcs.AI2025-05被引 3

算法创新可在算力受限下推动大模型进步,但效果因方法而异。

Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress

  • 区分算力依赖与独立的算法创新,用等效算力增益量化效果
  • 算力独立创新在各规模下提升性能,最高达3.5倍效率增益
  • 监管应关注算法研发,不能只限硬件,尤其适合政策制定者

针对大语言模型(LLM)发展的监管多聚焦于限制高性能计算资源获取。本研究评估此类措施的有效性,探究在算力受限环境下,是否可通过算法创新实现能力提升。提出新框架,区分算力依赖型创新(高算力下收益显著)与算力独立型创新(跨算力规模提升效率)。采用等效算力增益(CEG)量化影响。使用nanoGPT模型验证表明,算力独立创新在各测试规模下均带来显著性能提升(联合CEG最高达3.5×);而算力依赖型创新在小规模下反而损害性能,随模型规模增大,其CEG趋于基线水平,符合其在高算力下才产生主要效益的定义。关键发现:算力限制虽可能延缓进展,却无法阻止由算法进步带来的能力提升。因此,有效人工智能治理需纳入对算法研究的理解、预测与引导机制,超越单一硬件管控。该框架亦可用于预测AI发展路径。

原文摘要 · Abstract (English)

Regulatory efforts to govern large language model (LLM) development have predominantly focused on restricting access to high-performance computational resources. This study evaluates the efficacy of such measures by examining whether LLM capabilities can advance through algorithmic innovation in compute-constrained environments. We propose a novel framework distinguishing compute-dependent innovations--which yield disproportionate benefits at high compute--from compute-independent innovations, which improve efficiency across compute scales. The impact is quantified using Compute-Equivalent Gain (CEG). Experimental validation with nanoGPT models confirms that compute-independent advancements yield significant performance gains (e.g., with combined CEG up to $3.5\times$) across the tested scales. In contrast, compute-dependent advancements were detrimental to performance at smaller experimental scales, but showed improved CEG (on par with the baseline) as model size increased, a trend consistent with their definition of yielding primary benefits at higher compute. Crucially, these findings indicate that restrictions on computational hardware, while potentially slowing LLM progress, are insufficient to prevent all capability gains driven by algorithmic advancements. We argue that effective AI oversight must therefore incorporate mechanisms for understanding, anticipating, and potentially guiding algorithmic research, moving beyond a singular focus on hardware. The proposed framework also serves as an analytical tool for forecasting AI progress.

大模型算法创新算力限制治理策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。