arXiv:2509.25219cs.ITcs.CV2025-09

用统一评分框架解决压缩算法选型难题,兼顾压缩率与速度。

Challenges and Solutions in Selecting Optimal Lossless Data Compression Algorithms

  • 构建融合压缩率、编码/解码时间的统一评分模型,合理加权
  • 实验验证框架能准确识别不同优先级下的最优压缩算法
  • 发现学习型算法压缩率高,经典算法在速度上仍占优

数字数据的快速增长加剧了对高效无损压缩方法的需求。然而,现有算法存在权衡:部分压缩率高,部分编码或解码速度快,尚无算法在所有维度均表现最优。这种不匹配使多指标关键场景(如医学影像)中的算法选择变得复杂,因该场景需兼顾紧凑存储与快速检索。为此,我们提出一种数学框架,将压缩比、编码时间和解码时间整合为统一性能评分。通过合理加权与归一化,实现对不同算法的客观公平比较。在图像与文本数据集上的大量实验验证了该方法的有效性,结果表明其能可靠识别不同优先级设置下的最优压缩器。此外,研究发现现代基于学习的编解码器通常提供更优压缩比,但当速度为首要考量时,经典算法仍具优势。该框架为选择最优无损数据压缩技术提供了稳健且可适应的决策支持工具,弥合了理论度量与实际应用需求之间的差距。

原文摘要 · Abstract (English)

The rapid growth of digital data has heightened the demand for efficient lossless compression methods. However, existing algorithms exhibit trade-offs: some achieve high compression ratios, others excel in encoding or decoding speed, and none consistently perform best across all dimensions. This mismatch complicates algorithm selection for applications where multiple performance metrics are simultaneously critical, such as medical imaging, which requires both compact storage and fast retrieval. To address this challenge, we present a mathematical framework that integrates compression ratio, encoding time, and decoding time into a unified performance score. The model normalizes and balances these metrics through a principled weighting scheme, enabling objective and fair comparisons among diverse algorithms. Extensive experiments on image and text datasets validate the approach, showing that it reliably identifies the most suitable compressor for different priority settings. Results also reveal that while modern learning-based codecs often provide superior compression ratios, classical algorithms remain advantageous when speed is paramount. The proposed framework offers a robust and adaptable decision-support tool for selecting optimal lossless data compression techniques, bridging theoretical measures with practical application needs.

无损压缩算法选择性能评估医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。