让胸部X光报告生成模型精准聚焦病灶区域,实现高清感知不增加计算负担。
Seeing What Matters: Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Report Generation

- 通过自适应分配高分辨率区域,只对可疑病灶进行精细分析。
- 在不增加视觉令牌数的前提下,支持1920x1920原图级感知,效率提升超10倍。
- 适合需要高精度病灶识别的临床报告生成场景。
尽管胸部X光(CXR)基础模型发展迅速,多数放射科报告生成(RRG)系统仍依赖大幅下采样的输入(如256x256),受限于预训练视觉编码器的固定视觉令牌预算,抑制了原图中细微但关键的临床线索。然而,实现高分辨率(high-res)感知面临挑战:直接拼接导致令牌爆炸,全局压缩则削弱细微病灶,降低诊断保真度。受放射科医生工作流程启发,我们提出莱帕克斯(LePaX)——首个在不增加视觉令牌数前提下实现高效高分辨率胸部X光感知(最高达1920x1920)的报告生成框架。LePaX将高分辨率感知建模为在固定令牌预算下的空间分辨率分配问题,引入两个核心组件:可学习的空间分辨率分配(LSRA),通过学习空间效用图自适应地将有限的高分辨率能力分配给诊断相关区域,实现从原图中靶向提取高分辨率图像块;全局-局部融合(GRF),通过空间对齐的分辨率回写机制,将高分辨率区域证据投影回全局特征网格,完成保令牌的区域到全局精炼,避免令牌膨胀。在多个CXR基准上的实验表明,LePaX在临床与语言指标上均持续提升,同时以超过10倍于原始高分辨率拼接的效率实现原图级感知。
原文摘要 · Abstract (English)
Despite rapid advances in chest X-ray (CXR) foundation models, most radiology report generation (RRG) systems still rely on heavily downsampled inputs (e.g., 256x256) due to the fixed visual token budgets of pretrained vision encoders, suppressing subtle yet clinically important cues present in native-resolution images. However, enabling high-resolution (high-res) perception remains challenging: naive tiling causes prohibitive token inflation, while global compression suppresses subtle lesions and degrades diagnostic fidelity. Inspired by radiologists' workflow, localizing suspicious regions before detailed high-res assessment. We propose Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Reporting (LePaX), the first RRG framework that enables efficient high-res CXR perception (up to 1920x1920) without increasing the vision-token count. LePaX formulates high-res perception as a constrained spatial resolution allocation problem under a fixed token budget and introduces two key components: Learnable Spatial Resolution Allocation (LSRA), which learns a spatial utility map that adaptively allocates limited high-res capacity to diagnostically relevant regions, enabling targeted extraction of high-res patches from native CXRs; and Global-Regional Fusion (GRF), which performs token-preserving region-to-global refinement by projecting high-resolution regional evidence back onto the global feature grid through spatially aligned resolution write-back, avoiding token inflation. Experiments on multiple CXR benchmarks demonstrate that LePaX consistently improves both clinical and linguistic metrics while enabling native-resolution CXR perception with over 10x fewer visual tokens than naive high-res tiling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。