arXiv:2601.10455cs.CLcs.RO2026-01

用目标达成度重新评估手术规划,更准确识别模型真伪。

SurgGoal: Rethinking Surgical Planning Evaluation via Goal-Satisfiability

  • 以手术阶段目标是否达成定义规划正确性,结合专家规则
  • 发现序列相似性指标会误判有效计划,且漏检无效计划
  • 强调结构知识重要性,适合医疗AI评估研究者参考

手术规划融合视觉感知、长程推理与操作知识,但现有评估协议在安全关键场景中是否可靠仍不明确。基于目标导向的手术规划视角,我们通过阶段目标达成度定义规划正确性,依据专家设定的手术规则判定计划有效性。在此基础上,构建包含有效流程变体与含顺序及内容错误的无效计划的多中心元评估基准。实验表明,序列相似性指标会系统性误判规划质量,既惩罚有效计划又无法识别无效计划。因此,采用基于规则的目标达成度指标作为高精度元评估参考,在逐步受限条件下评估视频大语言模型,揭示出因感知错误和约束不足导致的失败。结构知识始终提升性能,而仅依赖语义引导不可靠,仅当结合结构约束时,大模型才获益。

原文摘要 · Abstract (English)

Surgical planning integrates visual perception, long-horizon reasoning, and procedural knowledge, yet it remains unclear whether current evaluation protocols reliably assess vision-language models (VLMs) in safety-critical settings. Motivated by a goal-oriented view of surgical planning, we define planning correctness via phase-goal satisfiability, where plan validity is determined by expert-defined surgical rules. Based on this definition, we introduce a multicentric meta-evaluation benchmark with valid procedural variations and invalid plans containing order and content errors. Using this benchmark, we show that sequence similarity metrics systematically misjudge planning quality, penalizing valid plans while failing to identify invalid ones. We therefore adopt a rule-based goal-satisfiability metric as a high-precision meta-evaluation reference to assess Video-LLMs under progressively constrained settings, revealing failures due to perception errors and under-constrained reasoning. Structural knowledge consistently improves performance, whereas semantic guidance alone is unreliable and benefits larger models only when combined with structural constraints.

手术规划目标评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。