arXiv:2608.03890cs.CVcs.AI2026-08

让医学影像模型不仅会写报告,还能精准定位病灶并自动测量,临床可用性大幅提升。

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

论文配图:CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
图 1 · 摘自论文原文
  • 融合分类与定位监督,生成报告同时支持可调阈值诊断和空间定位。
  • 在四个基准上表现领先,VQA准确率达94.0%,空间解码接近专用检测头水平。
  • 引入工具调用机制,对5种依赖测量的疾病诊断,平均F1提升43.6个百分点。

一个临床实用的胸部X光系统不仅要能流畅生成报告,还应具备可调决策阈值的分类能力、精确的空间定位以及生成诊断所依赖的解剖测量。当前的视觉语言模型(VLM)大多将这些任务分开处理,甚至忽略,导致与放射科医生实际需求之间存在差距。我们提出CARE-X,一种整合辅助判别监督与奖励对齐生成的胸部X光VLM。CARE-X在其生成主干中加入焦点损失分类头和复合损失定位头,与语言建模目标联合训练,实现可调阈值的诊断预测与精确定位,同时提升报告质量,证明结构化预测与生成相互促进。在此基础上,采用解耦式Clip与动态采样策略优化(DAPO),利用任务特异性奖励信号直接优化临床相关指标,涵盖报告生成、视觉问答(VQA)与空间定位。结果在四个报告生成基准上取得多数指标最优,ReXVQA上达到94.0%的准确率(较次优基线+6.0个百分点),生成式空间解码接近专用检测头性能。此外,为应对依赖测量的诊断,我们将Qwen3-VL-4B-Instruct与原生工具调用能力结合,调用确定性测量工具,同时保持对图像的完整视觉访问。该混合推理方式在五个测量依赖性条件上,平均F1相比仅依赖感知的基线提升43.6个百分点。

原文摘要 · Abstract (English)

A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vision-Language Models (VLMs) treat these as separate problems, if they address them at all, leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained alongside the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality, providing evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, visual question answering (VQA), and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report-generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct with native tool-calling capabilities for invoking deterministic measurement tools, while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions.

医学影像视觉语言模型工具调用可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。