arXiv:2608.21140cs.CVcs.AI2026-08

用模块化设计提升CT影像空间关系判断的可靠性与可审计性

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

论文配图:A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
图 1 · 摘自论文原文
  • 分三步走:解析语言、定位器官、用几何规则验证空间关系
  • 在MIRP基准上达到94.1%准确率,比端到端模型高42.5个百分点
  • 适合需要可解释医疗报告生成的场景,增强诊断可信度

可靠的时空理解是未来支持放射科报告生成和结构化图像理解的医学视觉-语言系统的重要前提。尽管现代视觉-语言模型(VLMs)在诸多医学影像任务中表现良好,但其在受控空间推理方面仍显薄弱,常无法可靠地基于图像证据定位空间关系。由于放射科推理依赖于解剖结构与病灶的相对位置,这一缺陷可能影响诊断准确性。本文提出一种模块化医学影像代理,用于轴向CT切片中的二元空间关系验证。系统不直接端到端预测空间答案,而是分解为明确阶段:语言解析、解剖定位和确定性几何验证。自然语言查询被转化为结构化关系元组,目标器官通过基于YOLO的检测器定位,最终的空间决策由物体中心通过确定性几何规则计算得出。我们在保留的MIRP空间问答基准上评估该方法,并与代表性端到端VLM基线进行对比。最佳混合配置达到94.1%准确率和94.2% F1,比直接Qwen2-VL提示高出42.5个百分点准确率,同时保持可解释的中间表示和可审计的推理阶段。结果表明,显式模块化空间验证可作为未来面向报告的医学影像代理的有力构建模块。

原文摘要 · Abstract (English)

Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language models (VLMs) show promising performance on many medical imaging tasks, recent evidence suggests they remain weak in controlled spatial reasoning and often fail to reliably ground spatial relations in image evidence. Given that radiological reasoning hinges on understanding the relative positions of anatomical structures and findings, this spatial weakness poses risks to diagnostic accuracy. We present a modular medical imaging agent for binary spatial relation verification in axial CT slices. Instead of directly predicting spatial answers end-to-end, the system decomposes the task into explicit stages: language parsing, anatomical localization, and deterministic geometric verification. Natural-language queries are converted into structured relation tuples, queried organs are localized with a YOLO-based detector, and the final spatial decision is computed from object centers using deterministic geometric rules. We evaluate the approach on the held-out MIRP spatial QA benchmark and compare it against representative end-to-end VLM baselines. The best-performing hybrid configuration reaches 94.1% accuracy and 94.2% F1, outperforming direct Qwen2-VL prompting by 42.5 percentage points in accuracy, while preserving interpretable intermediate representations and auditable reasoning stages. The results suggest that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.

医学影像空间推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。