arXiv:2606.20852cs.CVcs.AI2026-06

不微调模型,用激活控制提升肺部X光肺炎诊断准确率

Translating Inference-Time Control to Radiology Vision-Language Models: Activation Steering for Pneumonia Classification on Chest X-rays

论文配图:Translating Inference-Time Control to Radiology Vision-Language Models: Activation Steering for Pneumonia Classification on Chest X-rays
图 1 · 摘自论文原文
  • 通过文本和图像对比向量在推理时调节模型激活
  • 最佳模型F1分数从0.7692提升至0.8727,显著改善诊断性能
  • 适合希望快速优化医疗视觉语言模型的临床研究者

推理时工程可在不微调模型的情况下改变其行为。然而,该方法在医学视觉语言模型(VLMs)中提升诊断性能的效果尚不明确。本研究评估了对比激活添加(CAA)是否能在不更新权重的前提下提升胸部X光片上肺炎分类的表现。使用三个冻结的胸部X光VLM(MedGemma-4B-IT、NV-Reason-CXR-3B、CheXOne-3B),在公开的Kermany肺炎测试集上进行评估。分类基于二元提示下‘是’与‘否’标记的逻辑值。调控向量包括30对答案偏差控制、30对肺炎文本对比,以及基于30张肺炎和30张正常发展图像生成的图像条件对比。使用100张图像的确定性开发集进行层与尺度选择,另100张用于阈值校准。通过ROC-AUC、PR-AUC、F1分数、阈值分析、反向向量控制、随机向量控制及条件自举置信区间评估性能。固定阈值下的F1提升频繁出现,但未持续反映诊断性能改善。对于MedGemma-4B-IT,NV-Reason-CXR-3B表现最佳:经肺炎文本调控后校准F1从零样本的0.7692升至0.8619,图像条件调控后达0.8727。对于CheXOne-3B,肺炎文本调控使校准F1从0.8528升至0.8666,但置信区间跨越零。在这一公开肺炎基准上,CAA在无微调情况下显著改变预测分数分布与操作特征。其中一模型实现有意义性能提升,表明激活操控可作为适应医学VLM行为的轻量级方法。

原文摘要 · Abstract (English)

Inference-time engineering can alter model behavior without fine-tuning. However, its utility for improving diagnostic performance in medical vision-language models (VLMs) remains unclear. We aim to evaluate whether Contrastive Activation Addition (CAA) can improve pneumonia classification in chest radiograph VLMs without updating model weights. Three frozen chest radiograph VLMs (MedGemma-4B-IT, NV-Reason-CXR-3B, and CheXOne-3B) were evaluated on the public Kermany pneumonia test set. Classification was based on the logits of the tokens Yes and No under a binary prompt. Steering vectors included a 30-pair answer-bias control, a 30-pair pneumonia text contrast, and an image-conditioned contrast derived from 30 pneumonia and 30 normal development images. A deterministic 200-image development set was used for layer and scale selection (100 images) and threshold calibration (100 images). Performance was assessed using ROC-AUC, PR-AUC, F1 score, threshold analyses, reverse-vector controls, random-vector controls, and conditional bootstrap confidence intervals. Fixed-threshold F1 improvements were frequently observed but did not consistently indicate improved diagnostic performance. For MedGemma-4B-IT. NV-Reason-CXR-3B showed the strongest benefit: calibrated F1 improved from 0.7692 in the zero-shot setting to 0.8619 with pneumonia-text steering and to 0.8727 with image-conditioned steering. For CheXOne-3B, pneumonia-text steering increased calibrated F1 from 0.8528 to 0.8666, although the confidence interval crossed zero. On this public pneumonia benchmark, CAA substantially altered prediction score distributions and operating characteristics without fine-tuning. Meaningful performance gains were observed in one of three evaluated VLMs, suggesting that activation steering may serve as a lightweight approach for adapting medical VLM behavior.

医学影像视觉语言模型推理调控肺炎检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。