用多模态数据+描述性提示,实时检测手术机器人操作错误
Real-Time Multimodal Activity-Aware Error Detection in Robot-Assisted Surgery

- 融合视频、动作轨迹和文本提示,识别手术中的细微错误
- 在JIGSAWS和SAR-RARP50数据集上分别提升5%和16.6%的准确率
- 适合关注手术安全与智能辅助系统的研究人员
机器人辅助微创手术虽提升精度,但操作复杂,技术性错误检测对患者安全至关重要。现有基于视频的执行错误检测方法常忽略手术流程中活动与错误类型的细粒度上下文信息,且未能充分融合多模态数据。本文提出统一框架,利用视频、运动学数据及描述性文本提示进行执行错误检测。通过活动提示机制,将手势级活动、器械-物体交互及错误定义融入语言描述。引入基于预训练手术活动标签的活动感知视觉嵌入,对比了对比语言-图像嵌入与传统图像嵌入在错误检测中的效果。通过无缝融合运动学数据与视频、文本模态,框架显著提升检测性能,在JIGSAWS和SAR-RARP50数据集上分别取得5%和16.6%的F1分数提升,验证了结构化文本提示与多模态数据结合的价值。
原文摘要 · Abstract (English)
Robot-assisted minimally invasive surgery improves surgical precision but introduces complexity, making technical error detection essential for ensuring patient safety. Current executional error detection methods using video data often overlook fine-grained contextual descriptions of activities and error types within the hierarchical structure of surgical procedures. They also under-utilize complementary multimodal information. We propose a unified framework for executional error detection that leverages multimodal input, including video, kinematics, and descriptive textual prompts. Through activity prompting, we integrate descriptive language in gesture-level activities, instrument-object interactions, and error definitions. We also introduce activity-aware visual embeddings derived from vision encoders pretrained on surgical activity labels to compare the effectiveness of contrastive language-image embeddings with traditional image-based embeddings for error detection. By seamlessly integrating kinematic data with video and textual modalities, our framework significantly improves error detection performance. Achieving up to 5\% and 16.6\% F1 score improvements over state-of-the-art baselines on the JIGSAWS and SAR-RARP50 datasets, respectively, we demonstrate the value of combining curated textual prompts with multimodal data for accurate error detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。