让心脏手术报告生成更靠谱,减少幻觉错误。
TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

- 引入风险驱动的因果对齐机制,确保文字与影像真实对应
- 在1482例患者数据上,幻觉率降至8.1%,指标达新高
- 适合医疗AI、影像诊断等需要高可靠性的场景
经导管主动脉瓣置换术(TAVR)规划需精准多模态推理。但将多模态大模型应用于此高风险领域时,常因诊断幻觉导致生成内容缺乏解剖学依据。为此提出TAVR-VLM:一种新型框架,包含风险条件因果对齐注意力(R-CGA),构建模型内部的“风险→区域→词”结构化对齐路径。R-CGA将多模态输入压缩为因果风险瓶颈,将密集视觉特征提炼为全局风险掩码。自回归生成时,通过支持投影的因果一致性目标,在风险定义的支持掩码内约束词级对齐。在包含1,482名患者的M³TAVR数据集上评估,TAVR-VLM达到新基准:AUROC为0.896,CIDEr提升至0.936,幻觉率降至8.1%,显著提升基于证据的外科AI可解释性。
原文摘要 · Abstract (English)
Transcatheter Aortic Valve Replacement (TAVR) planning requires meticulous multimodal reasoning. However, adapting Multimodal Large Language Models (MLLMs) to this high-stakes domain is severely impeded by diagnostic hallucinations, where generated text lacks anatomical grounding. To address this, TAVR-VLM is introduced: a novel framework featuring Risk-Conditioned Causal Grounding Attention (R-CGA) that instantiates a model-internal ``Risk $\rightarrow$ Region $\rightarrow$ Word'' structural grounding pathway. R-CGA compresses multimodal inputs into a causal risk bottleneck, purifying dense visual features into a global risk mask. During autoregressive generation, a support-projected causal consistency objective constrains token-level grounding within the risk-defined support mask. Evaluated on $\text{M}^3\text{TAVR}$, a comprehensive 1,482-patient cohort, TAVR-VLM establishes a new state-of-the-art. It achieves an AUROC of 0.896, boosts CIDEr to 0.936, and drastically reduces the hallucination rate to 8.1\%, thereby improving interpretability for evidence-based surgical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。