用多模态AI自动生成手术视频摘要,提升记录效率与准确性。
Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI
- 分三步:先提取视频帧特征,再生成描述,最后整合成报告。
- 工具检测精度达96%,时间上下文摘要BERT分数0.74。
- 适合医疗AI、手术训练与智能病历系统研究者参考。
自动总结手术视频对改进手术记录、支持外科培训和术后分析至关重要。本文提出一种融合人工智能与医学的新型多模态框架,利用计算机视觉与大语言模型生成全面的视频摘要。方法分为三阶段:首先将手术视频分割为片段,使用视觉变压器在帧级别提取工具、组织、器官和手术动作特征;其次,通过大语言模型将特征转换为帧级描述,并结合基于ViViT的时序特征编码器生成片段级摘要;最后,采用专用于摘要任务的大语言模型聚合片段描述,生成完整手术报告。在包含50例腹腔镜视频的CholecT50数据集上评估,工具检测精度达96%,时序上下文摘要的BERTScore为0.74。该工作推动了辅助手术报告的AI工具发展,迈向更智能可靠的临床文档化。
原文摘要 · Abstract (English)
The automatic summarization of surgical videos is essential for enhancing procedural documentation, supporting surgical training, and facilitating post-operative analysis. This paper presents a novel method at the intersection of artificial intelligence and medicine, aiming to develop machine learning models with direct real-world applications in surgical contexts. We propose a multi-modal framework that leverages recent advancements in computer vision and large language models to generate comprehensive video summaries. % The approach is structured in three key stages. First, surgical videos are divided into clips, and visual features are extracted at the frame level using visual transformers. This step focuses on detecting tools, tissues, organs, and surgical actions. Second, the extracted features are transformed into frame-level captions via large language models. These are then combined with temporal features, captured using a ViViT-based encoder, to produce clip-level summaries that reflect the broader context of each video segment. Finally, the clip-level descriptions are aggregated into a full surgical report using a dedicated LLM tailored for the summarization task. % We evaluate our method on the CholecT50 dataset, using instrument and action annotations from 50 laparoscopic videos. The results show strong performance, achieving 96\% precision in tool detection and a BERT score of 0.74 for temporal context summarization. This work contributes to the advancement of AI-assisted tools for surgical reporting, offering a step toward more intelligent and reliable clinical documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。