arXiv:2507.05011cs.AIcs.CV2025-07中稿 · the MICCAI2025 wor…被引 1

手术规划中模仿学习比强化学习更有效,因专家数据分布更优

DARIL: When Imitation Learning outperforms Reinforcement Learning in Surgical Action Planning

  • 用双重自回归框架实现模仿学习,精准预测手术动作三元组
  • 在CholecT50数据集上模仿学习达到34.6%的三元组识别准确率
  • 揭示了强化学习在真实手术场景中受限于专家数据分布的局限性

手术动作规划需实时预测未来器械-动词-目标三元组。尽管远程操控机器人手术可提供自然专家示范用于模仿学习(IL),强化学习(RL)则有望通过自我探索发现更优策略。我们在CholecT50数据集上首次全面比较了IL与RL在手术动作规划中的表现。提出的双任务自回归模仿学习(DARIL)基线在动作三元组识别上达到34.6% mAP,下一帧预测达33.6% mAP,且在10秒时间窗下仅轻微下降至29.2%。我们评估了三种RL变体:基于世界模型的RL、直接视频RL和逆向RL增强。令人意外的是,所有RL方法均逊于DARIL——世界模型RL在10秒时降至3.1% mAP,直接视频RL仅达15.9%。分析表明,在专家标注测试集上的分布匹配机制系统性偏爱模仿学习,即使某些强化学习策略可能更优。这一发现挑战了强化学习在序列决策中必然优越的假设,为手术AI发展提供了关键洞见。

原文摘要 · Abstract (English)

Surgical action planning requires predicting future instrument-verb-target triplets for real-time assistance. While teleoperated robotic surgery provides natural expert demonstrations for imitation learning (IL), reinforcement learning (RL) could potentially discover superior strategies through self-exploration. We present the first comprehensive comparison of IL versus RL for surgical action planning on CholecT50. Our Dual-task Autoregressive Imitation Learning (DARIL) baseline achieves 34.6% action triplet recognition mAP and 33.6% next frame prediction mAP with smooth planning degradation to 29.2% at 10-second horizons. We evaluated three RL variants: world model-based RL, direct video RL, and inverse RL enhancement. Surprisingly, all RL approaches underperformed DARIL--world model RL dropped to 3.1% mAP at 10s while direct video RL achieved only 15.9%. Our analysis reveals that distribution matching on expert-annotated test sets systematically favors IL over potentially valid RL policies that differ from training demonstrations. This challenges assumptions about RL superiority in sequential decision making and provides crucial insights for surgical AI development.

手术规划模仿学习强化学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。