arXiv:2505.14980eess.IV2025-05被引 6

为机器视觉编码设定理论极限,发现现有方法差距巨大。

Rate-Accuracy Bounds in Visual Coding for Machines

  • 基于信息论推导机器视觉编码的速率-精度理论边界。
  • 实测结果表明当前方法需高10到1000倍码率才能达理论精度。
  • 适合研究高效机器视觉压缩的学者与工程师参考。

越来越多的视觉信号(如图像、视频、点云)被采集仅用于计算机视觉模型的自动化分析,应用涵盖交通监控、机器人、自动驾驶、智能家居等。这一趋势催生了面向分析而非重建的压缩策略需求,即“机器编码”。本文借鉴离散无记忆信源有损编码理论,推导出若干典型机器视觉编码问题的速率-精度边界,并与文献中现有成果对比。结果显示,当前方法在达到特定精度时所需码率至少高出理论边界一个数量级,部分场景甚至高两到三个数量级。这表明机器视觉编码领域仍有巨大优化空间。

原文摘要 · Abstract (English)

Increasingly, visual signals such as images, videos and point clouds are being captured solely for the purpose of automated analysis by computer vision models. Applications include traffic monitoring, robotics, autonomous driving, smart home, and many others. This trend has led to the need to develop compression strategies for these signals for the purpose of analysis rather than reconstruction, an area often referred to as "coding for machines." By drawing parallels with lossy coding of a discrete memoryless source, in this paper we derive rate-accuracy bounds on several popular problems in visual coding for machines, and compare these with state-of-the-art results from the literature. The comparison shows that the current results are at least an order of magnitude -- and in some cases two or three orders of magnitude -- away from the theoretical bounds in terms of the bitrate needed to achieve a certain level of accuracy. This, in turn, means that there is much room for improvement in the current methods for visual coding for machines.

机器编码速率-精度信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。