用上下文感知提示提升CT图像质量评估,更真实可靠。
CAP-IQA: Context-Aware Prompt-Guided CT Image Quality Assessment
- 结合文本先验与图像上下文提示,分离理想标准与真实退化。
- 在LDCTIQA挑战赛中相关性得分2.8590,领先榜首4.24%。
- 适用于临床真实场景,尤其对儿科患者图像有强泛化能力。
基于提示的方法通过描述性文本嵌入医学先验知识,但在CT图像质量评估中应用极少。这类提示常引入偏差,因理想化定义未必适用于真实退化情况,如噪声、运动伪影或扫描仪差异。为此,我们提出上下文感知提示引导的图像质量评估框架(CAP-IQA),融合文本级先验与实例级上下文提示,并采用因果去偏技术,将理想化知识与图像特定退化分离开。该框架结合基于CNN的视觉编码器与领域专用文本编码器,评估腹部CT图像的诊断可见性、解剖清晰度和噪声感知。模型利用放射科风格提示与上下文感知融合,对齐语义与感知表示。在2023年低剂量CT图像质量评估挑战赛基准上,CAP-IQA总体相关性得分为2.8590(PLCC、SROCC与KROCC之和),超越最高排名团队的2.7427,提升4.24%。全面消融实验表明,提示引导融合与简化编码器设计共同提升了特征对齐与可解释性。此外,在包含91,514例儿科CT图像的自建数据集上的评估,验证了其在不同人群中的真实泛化能力。
原文摘要 · Abstract (English)
Prompt-based methods, which encode medical priors through descriptive text, have been only minimally explored for CT Image Quality Assessment (IQA). While such prompts can embed prior knowledge about diagnostic quality, they often introduce bias by reflecting idealized definitions that may not hold under real-world degradations such as noise, motion artifacts, or scanner variability. To address this, we propose the Context-Aware Prompt-guided Image Quality Assessment (CAP-IQA) framework, which integrates text-level priors with instance-level context prompts and applies causal debiasing to separate idealized knowledge from factual, image-specific degradations. Our framework combines a CNN-based visual encoder with a domain-specific text encoder to assess diagnostic visibility, anatomical clarity, and noise perception in abdominal CT images. The model leverages radiology-style prompts and context-aware fusion to align semantic and perceptual representations. On the 2023 LDCTIQA challenge benchmark, CAP-IQA achieves an overall correlation score of 2.8590 (sum of PLCC, SROCC, and KROCC), surpassing the top-ranked leaderboard team (2.7427) by 4.24%. Moreover, our comprehensive ablation experiments confirm that prompt-guided fusion and the simplified encoder-only design jointly enhance feature alignment and interpretability. Furthermore, evaluation on an in-house dataset of 91,514 pediatric CT images demonstrates the true generalizability of CAP-IQA in assessing perceptual fidelity in a different patient population.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。