arXiv:2505.19147cs.CLcs.AI2025-05被引 39

AI效率重心从模型压缩转向数据压缩,提升长上下文处理能力。

Shifting AI Efficiency From Model-Centric to Data-Centric Compression

  • 通过直接压缩训练和推理中的数据量来提升效率
  • 解决超长文本、高分辨率图像等场景下的自注意力二次计算瓶颈
  • 适合关注长序列建模与算力优化的研究者

大语言模型(LLMs)和多模态大语言模型(MLLMs)的发展长期依赖于模型参数的扩展。然而,随着硬件限制阻碍了模型进一步增长,主要计算瓶颈已转变为自注意力机制在日益增长的长序列(如超长文本、高分辨率图像和长视频)上的二次方开销。本文主张,高效人工智能的研究焦点正从模型为中心的压缩转向以数据为中心的压缩。我们提出数据为中心的压缩是新兴范式,通过直接压缩训练或推理过程中处理的数据量来提升效率。为此,我们建立了一个统一框架,形式化现有高效策略,并论证其对长上下文AI而言是一次关键范式转变。随后系统梳理了数据为中心压缩方法的现状,分析其在不同场景下的优势。最后,指出关键挑战与未来研究方向。本工作旨在提供新的视角,整合现有成果,并推动应对不断增长的上下文长度挑战的创新。

原文摘要 · Abstract (English)

The advancement of large language models (LLMs) and multi-modal LLMs (MLLMs) has historically relied on scaling model parameters. However, as hardware limits constrain further model growth, the primary computational bottleneck has shifted to the quadratic cost of self-attention over increasingly long sequences by ultra-long text contexts, high-resolution images, and extended videos. In this position paper, \textbf{we argue that the focus of research for efficient artificial intelligence (AI) is shifting from model-centric compression to data-centric compression}. We position data-centric compression as the emerging paradigm, which improves AI efficiency by directly compressing the volume of data processed during model training or inference. To formalize this shift, we establish a unified framework for existing efficiency strategies and demonstrate why it constitutes a crucial paradigm change for long-context AI. We then systematically review the landscape of data-centric compression methods, analyzing their benefits across diverse scenarios. Finally, we outline key challenges and promising future research directions. Our work aims to provide a novel perspective on AI efficiency, synthesize existing efforts, and catalyze innovation to address the challenges posed by ever-increasing context lengths.

AI效率数据压缩长上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。