arXiv:2503.18296cs.CL2025-03被引 10

用大模型预测手术下一步动作,提升智能手术辅助能力

Surgical Action Planning with Large Language Models

  • 基于大模型理解手术视频,预测未来操作步骤
  • 在CholecT50-SAP数据集上准确率达19.3%提升(微调后)
  • 适合手术辅助、教学与术中决策系统研发者

在机器人辅助微创手术中,我们提出手术动作规划(SAP)任务,旨在从视觉输入生成未来动作计划,以弥补当前智能系统缺乏术中预测性规划的不足。该任务有望增强术中指导并实现流程自动化,但面临理解器械-动作关系和追踪手术进程等挑战。大语言模型(LLMs)虽具备理解手术视频内容的潜力,却尚未被充分用于SAP中的预测决策,因现有研究多集中于回顾性分析。数据隐私、计算开销及模态特异性限制进一步凸显研究空白。为此,我们提出基于大模型的手术动作规划框架LLM-SAP,通过解析自然语言提示的手术目标,预测未来动作并生成文本响应,可用于手术教学、术中决策、流程记录与技能评估。该框架集成两个新模块:近历史聚焦记忆模块(NHF-MM)用于建模历史状态,以及动作规划提示生成器。我们在自建的CholecT50-SAP数据集上,使用Qwen2.5和Qwen2-VL模型进行评估,验证了其在下一步动作预测上的有效性。实验采用零样本设置测试预训练模型,并实施基于LoRA的监督微调(SFT)。结果显示,Qwen2.5-72B-SFT相比Qwen2.5-72B准确率提升19.3%。

原文摘要 · Abstract (English)

In robot-assisted minimally invasive surgery, we introduce the Surgical Action Planning (SAP) task, which generates future action plans from visual inputs to address the absence of intraoperative predictive planning in current intelligent applications. SAP shows great potential for enhancing intraoperative guidance and automating procedures. However, it faces challenges such as understanding instrument-action relationships and tracking surgical progress. Large Language Models (LLMs) show promise in understanding surgical video content but remain underexplored for predictive decision-making in SAP, as they focus mainly on retrospective analysis. Challenges like data privacy, computational demands, and modality-specific constraints further highlight significant research gaps. To tackle these challenges, we introduce LLM-SAP, a Large Language Models-based Surgical Action Planning framework that predicts future actions and generates text responses by interpreting natural language prompts of surgical goals. The text responses potentially support surgical education, intraoperative decision-making, procedure documentation, and skill analysis. LLM-SAP integrates two novel modules: the Near-History Focus Memory Module (NHF-MM) for modeling historical states and the prompts factory for action planning. We evaluate LLM-SAP on our constructed CholecT50-SAP dataset using models like Qwen2.5 and Qwen2-VL, demonstrating its effectiveness in next-action prediction. Pre-trained LLMs are tested in a zero-shot setting, and supervised fine-tuning (SFT) with LoRA is implemented. Our experiments show that Qwen2.5-72B-SFT surpasses Qwen2.5-72B with a 19.3% higher accuracy.

手术机器人大模型应用动作预测智能医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。