评测医疗AI在完整影像流程中生成可审计证据的能力
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
- 构建全流程医学影像评估框架MedFlowBench与可复现运行环境MedOpenClaw
- 多模型在需提供正确证据时性能大幅下降,暴露答案可信度问题
- 适合关注临床落地、可解释性与端到端医疗AI评估的研究者
医学影像评测常仅针对预选的2D图像、切片或区域进行,接近视觉识别任务。真实临床工作流则要求医生查阅完整影像研究,操作影像软件,跨切片和缩放级别导航,并提交可审计的视觉证据。本文认为,这一证据生成流程是当前医学影像代理评估中缺失的关键维度。为此,提出MedFlowBench——一个针对全研究级影像的VLM代理评估基准,以及MedOpenClaw——支持3D Slicer和QuPath等工具的可控可重放运行环境。每个评估回合中,代理需分析完整放射科研究或全幻灯片病理图像,返回任务答案并提交结构化证据(如关键切片、坐标、感兴趣区或病灶状态字段),这些证据将自动与隐藏的掩码、标注和标签比对。实验发现,仅评分最终答案会导致严重高估性能;当答案必须由正确证据支撑时,复杂流程中的表现显著下降。进一步表明,单纯增加图像分析工具无法解决问题:工具仅在能简化并可靠执行复杂步骤时才有效,而代理仍难以在多步操作中自主选择输入、管理视图状态并验证中间输出。MedFlowBench揭示了医疗影像代理能否从完整研究中生成可审计证据,而非仅从选定图像生成看似合理的答案。
原文摘要 · Abstract (English)
Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition. Real clinical workflows impose a different burden: readers must search through complete studies, operate imaging software, navigate across slices and magnifications, and document visual evidence that can be audited. We argue that this evidence-producing workflow is a critical missing evaluation axis for medical imaging agents. To study it, we introduce MedFlowBench, a full-study benchmark for VLM agents, together with MedOpenClaw, a controlled and replayable runtime in which agents operate medical imaging viewers such as 3D Slicer and QuPath. In each episode, an agent inspects a complete radiology study or whole-slide pathology image, returns a task answer, and submits structured evidence, including key slices, coordinates, regions of interest, or lesion-state fields. This evidence is automatically checked against withheld masks, annotations, and labels. Across evaluated models, final answer-only scoring gives an overly optimistic picture: when answers must also be supported by correct evidence, performance drops substantially on complex workflows. We further find that adding image-analysis tools does not by itself solve the problem. Tools help when they make a complex procedure simple and reliable, but agents still struggle when they must choose inputs, manage viewer state, and verify intermediate outputs over multiple steps. MedFlowBench exposes whether medical imaging agents can produce auditable evidence from complete studies, rather than plausible answers from selected images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。