arXiv:2604.09037cs.CVcs.CL2026-04被引 1

首个评估医学操作视频中流程状态演进的基准,揭示大模型在连续动作判断上的严重不足。

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos

  • 构建基于评分标准的全流程动作状态追踪评估框架
  • 多模型在中间步骤判断上表现差,即使整体流程评分看似合理
  • 核心瓶颈在于理解连续交互如何动态更新操作状态

当前多模态大模型视频基准主要关注事件识别、时间排序和长上下文回忆,却忽视了专家级流程判断的关键能力:跟踪正在进行的交互如何更新流程状态,从而决定后续动作的正确性。我们提出 SiMing-Bench,首个从完整临床技能视频中评估该能力的基准。它聚焦于基于评分标准的流程级判断,检验交互驱动的状态更新是否在整套操作中保持正确性。该基准基于 SiMing-Score,包含由医生标注的真实临床技能考试视频,涵盖心肺复苏、自动体外除颤器操作和袋阀面罩通气,每段视频均配有标准化分步评分表和双专家标注。在多种开源与闭源多模态大模型上测试发现,其与医生判断的一致性普遍偏低。即使整体流程相关性看似可接受,中间步骤的判别仍表现不佳,表明粗粒度全局评估严重高估了现有模型的流程判断能力。进一步的二分类步骤判断与对齐片段分析表明,瓶颈并非仅在于细粒度打分或时间定位,而在于建模连续交互如何随时间推移动态更新流程状态。

原文摘要 · Abstract (English)

Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but overlook a harder capability required for expert procedural judgment: tracking how ongoing interactions update the procedural state and thereby determine the correctness of later actions. We introduce SiMing-Bench, the first benchmark for evaluating this capability from full-length clinical skill videos. It targets rubric-grounded process-level judgment of whether interaction-driven state updates preserve procedural correctness across an entire workflow. SiMing-Bench is instantiated with SiMing-Score, a physician-annotated dataset of real clinical skill examination videos spanning cardiopulmonary resuscitation, automated external defibrillator operation, and bag-mask ventilation, each paired with a standardized step-wise rubric and dual-expert labels. Across diverse open- and closed-source MLLMs, we observe consistently weak agreement with physician judgments. Moreover, weak performance on rubric-defined intermediate steps persists even when overall procedure-level correlation appears acceptable, suggesting that coarse global assessment substantially overestimates current models' procedural judgment ability. Additional analyses with binary step judgment and step-aligned clips indicate that the bottleneck is not merely fine-grained scoring or temporal localization, but modeling how continuous interactions update procedural state over time.

医疗视频流程判断多模态模型状态追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。