arXiv:2605.00448cs.CVeess.IV2026-05

用压缩CT训练医疗影像模型,效率提升近半且精度损失小于7%。

Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis

论文配图:Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
图 1 · 摘自论文原文
  • 通过特征注意力风格迁移,从高清图像学习结构关系并适配压缩输入
  • 在三个数据集上压缩输入下仍保持93%-97%的基准模型性能
  • 适合资源受限场景下的临床AI部署,如移动端或低带宽环境

人工智能在医学影像中的应用受限于体数据的高计算复杂度和资源消耗。尽管胸部CT体积数据比投影放射提供更多诊断信息,但其处理通常需高算力,且以NIfTI或DICOM格式存储的未压缩图像带来巨大负担。为应对低资源部署与高效电子传输需求,本文探索使用JPEG压缩的胸部CT进行胸腔异常检测。提出特征注意力风格迁移(FAST)框架,将高保真CT表示中的激活模式与结构关系迁移到基于压缩输入的时空视觉编码器中。结合基于格拉姆矩阵的注意力风格保留与双注意力特征对齐,实现对退化图像的鲁棒特征提取。进一步引入结构化因子投影(SFP),采用块张量列车分解替代密集投影层,使投影头参数减少近一半。对比学习管道CT-Lite整合上述组件,并采用基于SigLIP的多模态对齐目标。在CT-RATE、NIDCH和Rad-ChestCT数据集上的实验表明,尽管输入为压缩图像且参数显著减少,CT-Lite在所有数据集上的AUROC均保持在未压缩基线的5-7%以内,为资源受限条件下的AI临床评估开辟新路径。

原文摘要 · Abstract (English)

The deployment of artificial intelligence in medical imaging is hindered by high computational complexity and resource-intensive processing of volumetric data. Although chest computed tomography (CT) volumes offer richer diagnostic information than projection radiography, their use in AI-based diagnosis remains limited due to the computational burden of processing uncompressed volumetric images (typically stored in NIfTI or DICOM format). Addressing the growing need for low-resource deployment and efficient electronic data transfer, we investigate the utilization of JPEG-compressed chest CT volumes for thoracic abnormality detection. We propose Feature Attention Style Transfer (FAST), a novel distillation framework that transfers both activation patterns and structural relationships from high-fidelity CT representations to a spatiotemporal visual encoder operating on compressed inputs. By combining Gram-matrix-based attention style preservation with dual-attention feature alignment, FAST enables robust feature extraction from degraded volumes. Furthermore, we introduce Structured Factorized Projection (SFP), leveraging Block Tensor Train decomposition as a parameter-efficient alternative to dense projection layers, reducing projection-head parameters by almost half. Our contrastive learning pipeline, CT-Lite, integrates these components with a SigLIP-based multimodal alignment objective. Experiments on CT-RATE, NIDCH, and Rad-ChestCT demonstrate that CT-Lite achieves AUROC within 5-7\% of the uncompressed-input baseline across all three datasets, despite operating on compressed inputs with significantly fewer parameters, paving the way for AI-based clinical evaluation under resource constraints.

医学影像压缩感知轻量化模型特征迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。