无需训练,提升医学模型对微小病灶的识别能力。
EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

- 构建病理解剖原型库,对比可疑区域与正常模式。
- 通过反事实推理筛选病灶相关区域,避免增强正常组织。
- 基于形态引导残差增强,强化微小病灶在全局表征中的贡献。
医学视觉语言模型(VLM)在病变检测和报告生成中展现出潜力,但对微小病灶的敏感性不足,因其视觉线索稀疏、对比度低且嵌入复杂解剖背景。局部视觉标记在聚合过程中,弱病灶信号易被忽略,导致难以识别。现有方法多依赖领域预训练、临床术语引导对齐或可训练增强,需额外训练或模型适配,易过拟合特定病灶形态,限制其在冻结模型上的应用。为此,我们提出EasyLens,一种无需训练、即插即用的微小病灶表征增强方法。首先构建EasyBank,一个包含病灶原型与解剖感知正常参考的原型空间,用于对比可疑区域。EasyTag通过反事实原型推理选择病灶相关区域,避免增强正常组织。EasyAmplifier则通过形态引导残差增强,强化选定区域表示,提升其在全局图像嵌入中的贡献。在多个医学图像数据集及冻结的医学VLM主干上实验表明,EasyLens显著提升微小病灶检测性能,优于现有编码器增强基线。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generation. However, their practical utility remains limited by insufficient sensitivity to subtle lesions, whose visual evidence is often sparse, low-contrast, and embedded within complex anatomical context. As local visual tokens are aggregated, these weak lesion cues can become underrepresented in global image representations, making them difficult for medical VLMs to recognize. Existing efforts to improve lesion sensitivity mainly rely on medical-domain vision-encoder pre-training, clinical-term-guided alignment, or trainable pathological representation enhancement. Although effective, these approaches usually require additional training or model-specific adaptation and may overfit to particular disease morphologies, limiting their applicability to frozen medical VLMs. To address these limitations, we propose EasyLens, a training-free plug-and-play subtle-lesion representation amplifier for medical VLMs. EasyLens first constructs EasyBank, a pathology-anatomy prototype space that provides lesion-related prototypes and anatomy-aware normal references for comparing suspicious patches against both pathological and normal anatomical patterns. To avoid blindly amplifying normal tissues, EasyTag selects lesion-relevant patches through counterfactual prototype reasoning. To counteract the dilution of subtle lesion cues in global image representations, EasyAmplifier strengthens the selected lesion-relevant patch representations through morphology-guided residual enhancement, thereby increasing their contribution to the global image embedding. Experiments on multiple medical image datasets and frozen medical VLM backbones show that EasyLens improves subtle-lesion detection and outperforms existing encoder-enhancement baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。