构建医疗界面流程化视觉定位基准,评估模型在多步操作中的连续推理能力。
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
- 将医疗界面操作建模为多步骤结构化空间决策序列
- 包含216段任务视频与597个标注关键帧,每项任务含2-3个相互依赖步骤
- 提出严格顺序评估协议和六类失败分类,助力诊断模型临床可靠性
尽管多模态大模型发展迅速,其在高风险医疗软件环境中的可靠视觉定位能力仍缺乏深入研究。现有GUI基准大多聚焦孤立的单步定位查询,忽略了真实医疗界面中随流程演进、状态动态变化的多步协同操作需求。本文提出MedSPOT,一个面向临床GUI的流程感知序列定位基准。不同于以往将定位视为独立预测任务,MedSPOT将操作流程建模为一系列结构化空间决策。该基准包含216个任务驱动视频与597个标注关键帧,每个任务包含2至3个相互依赖的定位步骤,涵盖界面层级、上下文依赖与动态条件下的细粒度空间精度。为评估流程鲁棒性,提出严格顺序评估协议:首次定位错误即终止任务评估,明确测量多步流程中的错误传播。同时引入全面的失败分类体系,包括边缘偏差、小目标错误、无预测、近似误判、远距离误判及工具栏混淆,支持对模型在临床界面行为的系统性诊断。通过从孤立定位转向流程感知的序列推理评估,MedSPOT为医疗软件环境中多模态模型的评估建立了真实且安全关键的基准。代码与数据见:https://github.com/Tajamul21/MedSPOT。
原文摘要 · Abstract (English)
Despite the rapid progress of Multimodal Large Language Models (MLLMs), their ability to perform reliable visual grounding in high-stakes clinical software environments remains underexplored. Existing GUI benchmarks largely focus on isolated, single-step grounding queries, overlooking the sequential, workflow-driven reasoning required in real-world medical interfaces, where tasks evolve across independent steps and dynamic interface states. We introduce MedSPOT, a workflow-aware sequential grounding benchmark for clinical GUI environments. Unlike prior benchmarks that treat grounding as a standalone prediction task, MedSPOT models procedural interaction as a sequence of structured spatial decisions. The benchmark comprises 216 task-driven videos with 597 annotated keyframes, in which each task consists of 2 to 3 interdependent grounding steps within realistic medical workflows. This design captures interface hierarchies, contextual dependencies, and fine-grained spatial precision under evolving conditions. To evaluate procedural robustness, we propose a strict sequential evaluation protocol that terminates task assessment upon the first incorrect grounding prediction, explicitly measuring error propagation in multi-step workflows. We further introduce a comprehensive failure taxonomy, including edge bias, small-target errors, no prediction, near miss, far miss, and toolbar confusion, to enable systematic diagnosis of model behavior in clinical GUI settings. By shifting evaluation from isolated grounding to workflow-aware sequential reasoning, MedSPOT establishes a realistic and safety-critical benchmark for assessing multimodal models in medical software environments. Code and data are available at: https://github.com/Tajamul21/MedSPOT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。