arXiv:2609.06931cs.CV2026-09

CARDEA让冠脉造影诊断可审计,用空间证据提升AI可信度。

CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation

论文配图:CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation
图 1 · 摘自论文原文
  • 用三阶段训练让模型生成带框标注的可解释推理过程
  • 在跨域情况下诊断准确率达91%,媲美心脏科医生
  • 强化学习使报告生成能力显著提升,适合临床部署

侵入性冠状动脉造影(CAG)是诊断冠心病的金标准,但解读差异大。现有AI系统虽能提高一致性,却缺乏可审计的决策过程,且难以全面评估。我们开发了CARDEA,一个统一的大规模视觉-语言模型,作为CAG全流程分析的核心。它仅基于公开数据集和封闭式任务分三阶段训练:视觉特征对齐、自蒸馏链式框(CoB)冷启动,以及基于可验证奖励的强化学习(RLVR),鼓励推理中使用边界框。我们在两个研究级诊断任务上评估:主导性分类和复杂性评估,对比专用分类器和两位介入心脏病专家。报告生成未参与训练,零样本评估在外部队列上使用血管严重程度宏F1得分。CARDEA在分布内主导性分类略逊于分类器,但在分布外达到相当水平(准确率0.91 [95%置信区间, 0.86至0.95]),复杂性评估与心脏病专家相当(准确率0.90 [CI, 0.82至0.97])。仅经RLVR训练的模型提升了零样本报告生成能力,其血管严重程度宏F1达0.686 [CI, 0.664至0.707],高于基础模型(0.513)和始终正常的基线(0.312)两倍以上。CARDEA实现从原始多视角视频到关键帧选取再到研究级诊断的端到端流程,并暴露结论背后的空间证据。在可验证的封闭任务上进行强化学习,激发了监督模仿无法获得的开放式报告能力。临床应用需前瞻性验证其与专家心脏科医生的一致性。

原文摘要 · Abstract (English)

Invasive coronary angiography (CAG) is the gold standard for diagnosing coronary artery disease, but interpretation varies substantially among observers. Existing AI systems can improve consistency but lack auditable decision processes and are limited in comprehensive open-ended assessment, undermining clinician trust and clinical adoption readiness. We developed CARDEA, a unified large vision-language model that serves as the inference core of a CAG pipeline. It was trained solely on public datasets and closed-ended tasks in three stages: visual feature alignment, a self-distilled Chain-of-Box (CoB) cold start, and reinforcement learning with verifiable rewards (RLVR) with a CoB reward encouraging bounding-box use in the reasoning trace. We assessed its two study-level diagnoses, dominance classification and complexity assessment, against a dedicated classifier and two interventional cardiologists. Report generation was excluded from training and evaluated zero-shot across stages on an external cohort using vessel-severity macro-$F_1$. CARDEA trailed the classifier on in-distribution dominance but drew level under domain shift (accuracy, 0.91 [95% confidence interval (CI), 0.86 to 0.95]) and was comparable to the cardiologists on complexity assessment (accuracy, 0.90 [CI, 0.82 to 0.97]). Only RLVR improved zero-shot report generation, raising its vessel-severity macro-$F_1$ (0.686 [CI, 0.664 to 0.707]) above the untuned base model (0.513) and over twice the always-normal floor (0.312). CARDEA runs an end-to-end CAG pipeline from raw multi-view videos through keyframe selection to study-level diagnosis while exposing auditable spatial evidence behind its conclusions. RLVR on verifiable closed-ended tasks surfaced open-ended reporting ability that supervised imitation did not. Clinical use requires prospective validation against expert cardiologists.

医学影像可解释AI冠脉造影视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。