arXiv:2602.24111cs.CVcs.AI2026-02被引 4

用形式化验证确保医学影像模型推理逻辑正确

Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification

  • 将放射报告文本转化为可验证的逻辑命题,用数学工具检测推理漏洞
  • 在5个胸部X光数据集上发现模型存在保守判断和随机幻觉等隐藏错误
  • 适合关注医疗AI安全性的研究者与临床辅助系统开发者

视觉语言模型(VLMs)在生成放射科报告方面展现出潜力,但常出现逻辑不一致问题,如诊断结论与自身感知结果不符或遗漏应有推论。传统词汇度量对临床同义表达惩罚过重,且无法在无参考条件下捕捉此类推理缺陷。为此,我们提出一种神经符号验证框架,可确定性地审计VLM生成报告的内部一致性。该流程将自由文本放射学发现自动形式化为结构化命题,利用SMT求解器(Z3)和临床知识库,验证每个诊断主张是否被数学蕴含、属于幻觉或被遗漏。在五个胸部X光基准测试中评估七种VLM,验证器揭示了保守观察和随机幻觉等传统指标无法察觉的推理错误。在标注数据集上,基于求解器的蕴含约束可作为严格后验保障,系统消除不支持的幻觉,显著提升生成式临床助手的诊断严谨性与精确性。

原文摘要 · Abstract (English)

Vision-language models (VLMs) show promise in drafting radiology reports, yet they frequently suffer from logical inconsistencies, generating diagnostic impressions unsupported by their own perceptual findings or missing logically entailed conclusions. Standard lexical metrics heavily penalize clinical paraphrasing and fail to capture these deductive failures in reference-free settings. Toward guarantees for clinical reasoning, we introduce a neurosymbolic verification framework that deterministically audits the internal consistency of VLM-generated reports. Our pipeline autoformalizes free-text radiographic findings into structured propositional evidence, utilizing an SMT solver (Z3) and a clinical knowledge base to verify whether each diagnostic claim is mathematically entailed, hallucinated, or omitted. Evaluating seven VLMs across five chest X-ray benchmarks, our verifier exposes distinct reasoning failure modes, such as conservative observation and stochastic hallucination, that remain invisible to traditional metrics. On labeled datasets, enforcing solver-backed entailment acts as a rigorous post-hoc guarantee, systematically eliminating unsupported hallucinations to significantly increase diagnostic soundness and precision in generative clinical assistants.

医学AI逻辑验证生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。