arXiv:2507.06959cs.CVcs.AI2025-07被引 7

用反事实理由优化胸部X光视觉语言模型,减少幻觉并降低标注成本。

CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale

  • 通过置信度-相似度联合挖掘筛选难点样本,平衡偏好数据分布。
  • 仅用5%微调样本即提升8.93%性能,达到当前最优水平。
  • 无需额外专家标注,适合临床可解释性需求高的场景。

视觉语言模型在医疗应用中易产生幻觉,影响可靠性。尽管偏好优化可通过临床反馈缓解此问题,但面临训练样本临床无关、数据分布不均及专家标注成本高昂等挑战。为此,我们提出CheXPO,一种结合置信度-相似度联合挖掘与反事实理由的胸部X光偏好优化方法。首先,构建涵盖多种问题类型的细粒度多任务胸部X光视觉指令数据集用于监督微调(SFT)。随后,通过令牌级置信度分析识别SFT失败的难点样本,并利用相似性检索扩展这些样本以平衡偏好样本分布;合成的反事实理由提供细粒度临床偏好,无需额外专家输入。实验表明,仅使用5%的SFT样本,CheXPO即实现8.93%的相对性能提升,在多样化临床任务中达到最先进水平,为真实世界放射科应用提供可扩展且可解释的解决方案。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are prone to hallucinations that critically compromise reliability in medical applications. While preference optimization can mitigate these hallucinations through clinical feedback, its implementation faces challenges such as clinically irrelevant training samples, imbalanced data distributions, and prohibitive expert annotation costs. To address these challenges, we introduce CheXPO, a Chest X-ray Preference Optimization strategy that combines confidence-similarity joint mining with counterfactual rationale. Our approach begins by synthesizing a unified, fine-grained multi-task chest X-ray visual instruction dataset across different question types for supervised fine-tuning (SFT). We then identify hard examples through token-level confidence analysis of SFT failures and use similarity-based retrieval to expand hard examples for balancing preference sample distributions, while synthetic counterfactual rationales provide fine-grained clinical preferences, eliminating the need for additional expert input. Experiments show that CheXPO achieves 8.93% relative performance gain using only 5% of SFT samples, reaching state-of-the-art performance across diverse clinical tasks and providing a scalable, interpretable solution for real-world radiology applications.

视觉语言模型医学影像偏好优化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。