用反事实推理提升医学视频诊断准确率
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis

- 通过扩散模型生成假设病理下的组织演变过程
- 在宫颈镜与肠镜数据上提升2.6%至10.2%性能
- 适合临床决策支持系统研究者使用
医学视频诊断需从检查过程中动态组织反应推断临床决策。现有方法依赖端到端学习,存在重外观轻病理、缺乏临床先验、仅基于观测无反事实比较等问题。本文提出MedVCR框架,模拟临床诊断思维:包含基于扩散的反事实生成器,用于合成特定病理状态下的组织演化;反事实表征学习模块,通过时间一致性、病灶可分性与反事实对齐等临床规则编码诊断知识;双路径诊断策略,融合视频级评估与帧级反事实分析。在全监督(如宫颈镜)和弱监督(如肠镜)场景下,相比领先基线提升2.6%~10.2%。全面消融实验验证各组件有效性。代码将公开。
原文摘要 · Abstract (English)
Medical video diagnosis involves inferring clinical decisions from dynamic tissue responses throughout examination processes. Existing methods rely on an end-to-end learning paradigm that i) focuses on appearance rather than pathology, ii) lacks clinical priors, and iii) reasons solely from observations without counterfactual comparison. This work introduces MedVCR, a counterfactual reasoning framework that mimics clinical diagnostic thinking. MedVCR comprises three components: a Counterfactual Generator that synthesizes tissue evolution under specified pathological states via a diffusion-based manner; a Counterfactual Representation Learning module that encodes diagnostic knowledge through clinical rules (i.e., temporal consistency, pathological separability, and counterfactual alignment); and a Dual Diagnostic Prediction strategy that integrates video-level assessment with frame-level counterfactual analysis. MedVCR is evaluated under both fully supervised (e.g., colposcopy) and weakly supervised (e.g., colonoscopy) video diagnosis settings, yielding 2.6%-10.2% performance gains compared with leading baselines. Comprehensive ablation studies further validate the effectiveness of each component. The code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。