神经压缩导致图像语义失真,这篇论文提出分类框架帮助识别风险。
A Taxonomy of Miscompressions: Preparing Image Forensics for Neural Compression
- 提出三类图像压缩错误类型与高影响标记
- 发现压缩后视觉完美但语义已变,难以检测
- 适合图像取证、压缩安全研究者参考
神经压缩有望革新有损图像压缩。基于生成模型的最新方案在保持高感知质量的同时实现前所未有的压缩率,但牺牲了语义保真度。解压后的图像细节看似光学完美,但语义上已不同于原始图像,使压缩错误难以甚至无法检测。我们探索问题空间,提出一个初步的误压缩分类体系。该体系定义了三类‘发生什么’的误压缩类型,并引入二元‘高影响’标志,用于标识改变语义符号的误压缩。我们讨论该分类体系如何促进风险沟通及缓解方法的研究。
原文摘要 · Abstract (English)
Neural compression has the potential to revolutionize lossy image compression. Based on generative models, recent schemes achieve unprecedented compression rates at high perceptual quality but compromise semantic fidelity. Details of decompressed images may appear optically flawless but semantically different from the originals, making compression errors difficult or impossible to detect. We explore the problem space and propose a provisional taxonomy of miscompressions. It defines three types of 'what happens' and has a binary 'high impact' flag indicating miscompressions that alter symbols. We discuss how the taxonomy can facilitate risk communication and research into mitigations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。