arXiv:2508.09075cs.CV2025-08被引 4

将图像压缩模型扩大到10亿参数,发现性能随规模提升的规律。

Scaling Learned Image Compression Models up to 1 Billion

  • 基于HPCM模型,从6850万参数扩展至10亿参数
  • 发现测试损失与模型规模、训练算力呈幂律关系
  • 10亿参数模型实现当前最优率失真性能

近期大语言模型的发展揭示了智能与压缩之间存在强关联。学习型图像压缩作为现代数据压缩的核心任务,近年来取得显著进展。然而,现有模型规模受限,制约其表征能力,且模型规模对压缩性能的影响尚不明确。本文首次系统研究大规模学习型图像压缩模型的扩展,并通过拟合幂律关系揭示性能变化趋势。以最新SOTA模型HPCM为基线,将模型参数从6850万扩展至10亿,建立测试损失与模型规模、最优训练算力之间的幂律关系。结果表明存在可外推的规模增长趋势。实验显示,扩展后的HPCM-1B模型达到当前最优率失真性能。本工作希望激励未来对大规模压缩模型的探索,并深化压缩与智能间关系的研究。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) highlight a strong connection between intelligence and compression. Learned image compression, a fundamental task in modern data compression, has made significant progress in recent years. However, current models remain limited in scale, restricting their representation capacity, and how scaling model size influences compression performance remains unexplored. In this work, we present a pioneering study on scaling up learned image compression models and revealing the performance trends through scaling laws. Using the recent state-of-the-art HPCM model as baseline, we scale model parameters from 68.5 millions to 1 billion and fit power-law relations between test loss and key scaling variables, including model size and optimal training compute. The results reveal a scaling trend, enabling extrapolation to larger scale models. Experimental results demonstrate that the scaled-up HPCM-1B model achieves state-of-the-art rate-distortion performance. We hope this work inspires future exploration of large-scale compression models and deeper investigations into the connection between compression and intelligence.

图像压缩模型规模扩展规律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。