arXiv:2503.19495eess.IVcs.CV2025-03被引 1

让压缩图像仍能准确识别文字,适合低算力设备使用

End-to-End Semantic Preservation in Text-Aware Image Compression Systems

  • 压缩时专门保留文字特征,编码器计算量仅为OCR模块一半
  • 低比特率下文字识别准确率超越未压缩图像的OCR表现
  • 即使图像严重失真,也能通过模型恢复语义,适合机器理解

传统图像压缩侧重人眼视觉保真度,而面向机器的编码则关注自动化理解所需信息。本文提出端到端压缩框架,在保持文本特性的前提下实现高效压缩,编码器计算成本约为OCR模块的一半,适用于资源受限设备。当本地进行OCR不可行时,可对图像高效压缩并后续解码恢复文本内容。实验表明,在低比特率下文本提取准确率显著提升,甚至优于未压缩图像上的OCR表现。进一步研究发现,经过极端压缩的通用编码器仍具备保留隐含语义的能力,通过学习增强与识别模块,可从视觉退化的表示中恢复有意义信息。结果验证了在机器中心图像编码中,文本导向压缩与通用语义保全的可行性。

原文摘要 · Abstract (English)

Traditional image compression methods aim to reconstruct images for human perception, prioritizing visual fidelity over task relevance. In contrast, Coding for Machines focuses on preserving information essential for automated understanding. Building on this principle, we present an end-to-end compression framework that retains text-specific features for Optical Character Recognition (OCR). The encoder operates at roughly half the computational cost of the OCR module, making it suitable for resource-limited devices. When on-device OCR is infeasible, images can be efficiently compressed and later decoded to recover textual content. Experiments show significant improvements in text extraction accuracy at low bitrates, even outperforming OCR on uncompressed images. We further extend this study to general-purpose encoders, exploring their capacity to preserve hidden semantics under extreme compression. Instead of optimizing for visual fidelity, we examine whether compact, visually degraded representations can retain recoverable meaning through learned enhancement and recognition modules. Results demonstrate that semantic information can persist despite severe compression, bridging text-oriented compression and general-purpose semantic preservation in machine-centered image coding.

图像压缩OCR机器理解低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。