arXiv:2412.09231eess.IV2024-12

提出可同时支持人眼与机器视觉的医学影像编码框架。

Versatile Volumetric Medical Image Coding for Human-Machine Vision

  • 设计三维自编码器,学习切片间潜在特征以增强表达能力。
  • 编码后直接支持分割等机器分析,无需解码为像素。
  • 在重建质量与分割精度上均优于传统与神经方法。

神经图像压缩(NIC)因其在特征表示和数据优化方面的显著优势受到广泛关注。然而,现有针对体素医学影像的NIC方法大多仅关注提升人眼感知质量。这些方法需将数据解码回像素才能进行下游机器学习分析,降低了现代数字医疗中诊断与治疗的效率。本文提出一种适用于人眼与机器视觉的通用体素医学图像编码框架(VVMIC),使编码后的表示可直接用于多种分析任务,无需解码为像素。针对体素影像特有的三维结构,设计了通用体素自编码器(VVAE)模块,学习切片间的潜在表示,增强当前切片的表达能力,并生成用于后续重建与分割的中间解码特征。为进一步提升编码性能,构建了融合切片间潜在上下文、空间-通道上下文及层级超上下文的多维上下文模型。实验结果表明,该框架在保持高质量人眼重建的同时,实现了比多种传统与神经方法更优的机器视觉分割精度。

原文摘要 · Abstract (English)

Neural image compression (NIC) has received considerable attention due to its significant advantages in feature representation and data optimization. However, most existing NIC methods for volumetric medical images focus solely on improving human-oriented perception. For these methods, data need to be decoded back to pixels for downstream machine learning analytics, which is a process that lowers the efficiency of diagnosis and treatment in modern digital healthcare scenarios. In this paper, we propose a Versatile Volumetric Medical Image Coding (VVMIC) framework for both human and machine vision, enabling various analytics of coded representations directly without decoding them into pixels. Considering the specific three-dimensional structure distinguished from natural frame images, a Versatile Volumetric Autoencoder (VVAE) module is crafted to learn the inter-slice latent representations to enhance the expressiveness of the current-slice latent representations, and to produce intermediate decoding features for downstream reconstruction and segmentation tasks. To further improve coding performance, a multi-dimensional context model is assembled by aggregating the inter-slice latent context with the spatial-channel context and the hierarchical hypercontext. Experimental results show that our VVMIC framework maintains high-quality image reconstruction for human vision while achieving accurate segmentation results for machine-vision tasks compared to a number of reported traditional and neural methods.

医学影像神经压缩三维建模多模态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。