arXiv:2507.06140eess.IVcs.AI2025-07被引 4

用语言模型提升低剂量CT去噪,让图像更清晰、更可信。

LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models

  • 用视觉语言模型提取解剖语义,指导低剂量CT去噪
  • 在两个公开数据集上超越现有方法,细节保留更好
  • 可解释性强,适合医疗影像与AI结合研究者

低剂量计算机断层扫描(LDCT)虽降低辐射暴露,但常导致图像质量下降,影响诊断准确性。现有深度学习去噪方法多聚焦像素级映射,忽略高层语义引导。近年来视觉语言模型(VLMs)的发展表明,语言可有效捕捉结构化语义信息,为提升LDCT重建提供新路径。本文提出LangMamba,一种基于语言驱动的Mamba框架,利用预训练VLM提取的表示,增强正常剂量CT(NDCT)对低剂量图像的监督。该框架采用两阶段学习策略:首先预训练语言引导自编码器(LangAE),通过冻结的VLM将NDCT图像映射至富含解剖信息的语义空间;其次融合两个核心组件——语义增强高效去噪器(SEED),利用Mamba机制捕捉全局特征并强化局部语义;以及语言协同双空间对齐损失(LangDA),确保去噪后图像在感知与语义空间中均与NDCT一致。在两个公开数据集上的大量实验表明,LangMamba显著优于现有最先进方法,在细节保持与视觉保真度方面表现优异。值得注意的是,LangAE在未见数据集上展现强泛化能力,有效降低训练成本。此外,LangDA损失通过引入语言引导见解提升了可解释性,并支持即插即用。本研究揭示了语言作为监督信号在推进LDCT去噪中的潜力。代码已公开于https://github.com/hao1635/LangMamba。

原文摘要 · Abstract (English)

Low-dose computed tomography (LDCT) reduces radiation exposure but often degrades image quality, potentially compromising diagnostic accuracy. Existing deep learning-based denoising methods focus primarily on pixel-level mappings, overlooking the potential benefits of high-level semantic guidance. Recent advances in vision-language models (VLMs) suggest that language can serve as a powerful tool for capturing structured semantic information, offering new opportunities to improve LDCT reconstruction. In this paper, we introduce LangMamba, a Language-driven Mamba framework for LDCT denoising that leverages VLM-derived representations to enhance supervision from normal-dose CT (NDCT). LangMamba follows a two-stage learning strategy. First, we pre-train a Language-guided AutoEncoder (LangAE) that leverages frozen VLMs to map NDCT images into a semantic space enriched with anatomical information. Second, we synergize LangAE with two key components to guide LDCT denoising: Semantic-Enhanced Efficient Denoiser (SEED), which enhances NDCT-relevant local semantic while capturing global features with efficient Mamba mechanism, and Language-engaged Dual-space Alignment (LangDA) Loss, which ensures that denoised images align with NDCT in both perceptual and semantic spaces. Extensive experiments on two public datasets demonstrate that LangMamba outperforms conventional state-of-the-art methods, significantly improving detail preservation and visual fidelity. Remarkably, LangAE exhibits strong generalizability to unseen datasets, thereby reducing training costs. Furthermore, LangDA loss improves explainability by integrating language-guided insights into image reconstruction and offers a plug-and-play fashion. Our findings shed new light on the potential of language as a supervisory signal to advance LDCT denoising. The code is publicly available on https://github.com/hao1635/LangMamba.

低剂量CT视觉语言模型去噪可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。