解决胸部X光检索中复合临床查询匹配难题,提升精准率。
CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography

- 设计结构化基准,支持包含并列与否定的复杂临床查询检索。
- 在双病灶组合和否定查询上,Precision@5分别提升8.5和22.0个百分点。
- 适用于需要高精度临床影像检索的医生与研究者。
大型胸片档案难以搜索,因多数研究仅附自由文本报告而无结构化临床标注。视觉语言模型虽适配文本到图像检索,但现有生物医学模型主要优化报告到图像匹配,而非满足短临床查询。这导致目标错位:模型可能检索出包含查询词的图像,却无法满足完整临床约束,尤其对包含并列或否定的查询(如“肺不张且无肺炎”)表现不佳。我们提出CXR-Retrieve,一个面向组合式胸片文本到图像检索的结构化基准。该基准包含来自MIMIC-CXR-JPG官方测试集的5,159张测试图像,以及145个涵盖单个、并列及正负性表述的文本查询。相关性定义为检索图像是否满足所有声明的病理约束,而非是否匹配配对报告。我们进一步提出一种标签感知对比微调目标,使模型吸引具有兼容病理约束的图文对(包括共享确认的缺失),同时明确排斥矛盾对。基于域内CXR-CLIP检查点,本方法在双病理组合查询上将Precision@5提升8.5个百分点,在否定查询上提升22.0个百分点。结果表明,可靠的胸片检索需训练目标不仅捕捉提及的发现,还需建模其临床断言方式。
原文摘要 · Abstract (English)
Large chest radiography archives are difficult to search because most studies are paired only with free-text reports rather than structured clinical annotations. Vision-language models offer a natural interface for text-to-image retrieval, but current biomedical models are primarily optimized for report-to-image matching rather than for satisfying short clinical search queries. This creates an objective mismatch: a model may retrieve images related to words in the query while failing to satisfy the full clinical constraint, especially for conjunctions and negations such as ``atelectasis and no pneumonia.'' We introduce CXR-Retrieve, a structured benchmark for compositional chest X-ray text-to-image retrieval. The benchmark contains 5,159 test images from the official test-split of MIMIC-CXR-JPG and 145 textual queries spanning single and conjunction findings, both positive and negative. Relevance is defined by whether a retrieved image satisfies all asserted pathology constraints, rather than by whether it matches a paired report. We further propose a label-aware contrastive fine-tuning objective for clinical retrieval. Our method attracts image-text pairs with compatible asserted pathology constraints, including shared confirmed absences, while explicitly repelling contradictory pairs. Starting from the in-domain CXR-CLIP checkpoint, our method improves Precision@5 over CXR-CLIP by 8.5 percentage points on two-pathology conjunctions and by 22.0 percentage points on negation queries. These results show that reliable chest X-ray retrieval requires training objectives that model not only which findings are mentioned, but also how they are clinically asserted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。