首个评估医学操作视频中流程状态演进的基准,揭示大模型在连续动作判断上的严重不足。
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
- 构建基于评分标准的全流程动作状态追踪评估框架
- 多模型在中间步骤判断上表现差,即使整体流程评分看似合理
- 核心瓶颈在于理解连续交互如何动态更新操作状态
当前多模态大模型视频基准主要关注事件识别、时间排序和长上下文回忆,却忽视了专家级流程判断的关键能力:跟踪正在进行的交互如何更新流程状态,从而决定后续动作的正确性。我们提出 SiMing-Bench,首个从完整临床技能视频中评估该能力的基准。它聚焦于基于评分标准的流程级判断,检验交互驱动的状态更新是否在整套操作中保持正确性。该基准基于 SiMing-Score,包含由医生标注的真实临床技能考试视频,涵盖心肺复苏、自动体外除颤器操作和袋阀面罩通气,每段视频均配有标准化分步评分表和双专家标注。在多种开源与闭源多模态大模型上测试发现,其与医生判断的一致性普遍偏低。即使整体流程相关性看似可接受,中间步骤的判别仍表现不佳,表明粗粒度全局评估严重高估了现有模型的流程判断能力。进一步的二分类步骤判断与对齐片段分析表明,瓶颈并非仅在于细粒度打分或时间定位,而在于建模连续交互如何随时间推移动态更新流程状态。
原文摘要 · Abstract (English)
Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but overlook a harder capability required for expert procedural judgment: tracking how ongoing interactions update the procedural state and thereby determine the correctness of later actions. We introduce SiMing-Bench, the first benchmark for evaluating this capability from full-length clinical skill videos. It targets rubric-grounded process-level judgment of whether interaction-driven state updates preserve procedural correctness across an entire workflow. SiMing-Bench is instantiated with SiMing-Score, a physician-annotated dataset of real clinical skill examination videos spanning cardiopulmonary resuscitation, automated external defibrillator operation, and bag-mask ventilation, each paired with a standardized step-wise rubric and dual-expert labels. Across diverse open- and closed-source MLLMs, we observe consistently weak agreement with physician judgments. Moreover, weak performance on rubric-defined intermediate steps persists even when overall procedure-level correlation appears acceptable, suggesting that coarse global assessment substantially overestimates current models' procedural judgment ability. Additional analyses with binary step judgment and step-aligned clips indicate that the bottleneck is not merely fine-grained scoring or temporal localization, but modeling how continuous interactions update procedural state over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。