arXiv:2511.12256cs.CVcs.AI2025-11

用文本提示引导图像分析,提升低剂量CT图像质量评估精度

Prompt-Conditioned FiLM and Multi-Scale Fusion on MedSigLIP for Low-Dose CT Quality Assessment

  • 通过文本提示控制特征,让模型理解临床意图
  • 在1000张训练图上达到PLCC 0.9575,超越竞赛最优结果
  • 适合医疗影像质量评估、快速适配新任务的研究者

我们提出一种基于MedSigLIP的提示条件框架,通过特征逐元素线性调制(FiLM)和多尺度池化注入文本先验。文本提示引导局部图像块特征,实现数据高效学习与快速适应。架构通过独立回归头融合全局、局部与纹理感知池化,采用轻量级MLP进行整合,并以成对排序损失训练。在包含1000张训练图像的LDCTIQA2023公开数据集上,PLCC达0.9575,SROCC为0.9561,KROCC为0.8301,优于已发表的竞赛最佳方案,验证了提示引导方法的有效性。

原文摘要 · Abstract (English)

We propose a prompt-conditioned framework built on MedSigLIP that injects textual priors via Feature-wise Linear Modulation (FiLM) and multi-scale pooling. Text prompts condition patch-token features on clinical intent, enabling data-efficient learning and rapid adaptation. The architecture combines global, local, and texture-aware pooling through separate regression heads fused by a lightweight MLP, trained with pairwise ranking loss. Evaluated on the LDCTIQA2023 (a public LDCT quality assessment challenge) with 1,000 training images, we achieve PLCC = 0.9575, SROCC = 0.9561, and KROCC = 0.8301, surpassing the top-ranked published challenge submissions and demonstrating the effectiveness of our prompt-guided approach.

医学影像低剂量CT提示工程质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。