通过对比推理提升医学影像判读准确性,避免冗余证据干扰。
DoubleTake: Contrastive Reasoning for Faithful Decision-Making in Medical Imaging
- 构建对抗性证据集,平衡视觉相关性与多样性
- 在MediConfusion上准确率提升15%,误判显著减少
- 适合需要高可信决策的临床辅助诊断场景
医学影像中的精准决策依赖于对相似病状间细微差异的推理,但现有方法多基于最近邻检索,返回冗余证据并强化单一假设。本文提出一种对比式、文档感知的参考选择框架,利用ROCO嵌入和元数据,显式平衡视觉相关性、嵌入多样性与来源溯源性,构建更利于区分的紧凑证据集。针对ROCO虽提供大规模图文对但未明确参考选择策略的问题,我们发布可复现的选择协议与精调参考库,支持对比检索的系统研究。在此基础上,提出反事实-对比推理框架,通过有置信度的结构化成对比较,采用基于边距的决策规则聚合证据并实现可靠弃权。在MediConfusion基准测试中,该方法达到当前最优性能,相对之前方法提升近15%的集合级准确率,同时降低混淆并提高个体判读精度。
原文摘要 · Abstract (English)
Accurate decision making in medical imaging requires reasoning over subtle visual differences between confusable conditions, yet most existing approaches rely on nearest neighbor retrieval that returns redundant evidence and reinforces a single hypothesis. We introduce a contrastive, document-aware reference selection framework that constructs compact evidence sets optimized for discrimination rather than similarity by explicitly balancing visual relevance, embedding diversity, and source-level provenance using ROCO embeddings and metadata. While ROCO provides large-scale image-caption pairs, it does not specify how references should be selected for contrastive reasoning, and naive retrieval frequently yields near-duplicate figures from the same document. To address this gap, we release a reproducible reference selection protocol and curated reference bank that enable a systematic study of contrastive retrieval in medical image reasoning. Building on these contrastive evidence sets, we propose Counterfactual-Contrastive Inference, a confidence-aware reasoning framework that performs structured pairwise visual comparisons and aggregates evidence using margin-based decision rules with faithful abstention. On the MediConfusion benchmark, our approach achieves state-of-the-art performance, improving set-level accuracy by nearly 15% relative to prior methods while reducing confusion and improving individual accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。