arXiv:2602.13332cs.CVcs.AI2026-02被引 3

让AI像医生一样看手术视频,逐步定位关键画面做推理

MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling

  • 分步检索视频中关键片段,边思考边调用工具验证
  • 在多个医学视频任务上超越现有模型,跨域表现也领先
  • 适合医疗AI研究者和临床辅助系统开发者

长时序临床视频是基于视觉的临床决策核心,对手术机器人等应用日益重要。然而当前多模态大模型通常被动采样或弱定位地处理视频,难以迭代地定位、验证并解释预测结果。为此,我们提出MedScope,一种通过粗到精证据搜索进行临床视频推理的工具使用模型。它通过在中间推理阶段调用工具并验证获取的观察,生成更准确且明确锚定在时间局部化视觉证据上的预测。为解决高质量监督数据缺乏问题,我们构建了以证据为中心的细粒度临床视频数据集ClinVideoSuite。随后采用接地感知的组相对策略优化(GA-GRPO)训练MedScope,直接以对齐证据的奖励和加权优势强化工具使用。在完整与细粒度视频理解基准上,MedScope在领域内与跨域评估中均达到领先水平。本方法揭示了医疗AI代理通过工具集成推理实现真正‘以视频思考’的路径。代码、模型与数据将公开。

原文摘要 · Abstract (English)

Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models typically process videos with passive sampling or weakly grounded inspection, which limits their ability to iteratively locate, verify, and justify predictions with temporally targeted evidence. To close this gap, we propose MedScope, a tool-using clinical video reasoning model that performs coarse-to-fine evidence seeking over long-form procedures. By interleaving intermediate reasoning with targeted tool calls and verification on retrieved observations, MedScope produces more accurate and trustworthy predictions that are explicitly grounded in temporally localized visual evidence. To address the lack of high-fidelity supervision, we build ClinVideoSuite, an evidence-centric, fine-grained clinical video suite. We then optimize MedScope with Grounding-Aware Group Relative Policy Optimization (GA-GRPO), which directly reinforces tool use with grounding-aligned rewards and evidence-weighted advantages. On full and fine-grained video understanding benchmarks, MedScope achieves state-of-the-art performance in both in-domain and out-of-domain evaluations. Our approach illuminates a path toward medical AI agents that can genuinely "think with videos" through tool-integrated reasoning. We will release our code, models, and data.

医学AI视频推理工具调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。