用时序预测自动选最佳手术视角,提升多摄像头录像的清晰度和实用性。
TSP-OCS: A Time-Series Prediction for Optimal Camera Selection in Multi-Viewpoint Surgical Video Analysis
- 基于视觉与语义特征融合的时序预测模型,动态选择最优摄像机视角。
- 在6视角甲状腺手术视频上,长时预测准确率优于现有方法。
- 适合手术教学与医疗安全领域,可减少遮挡带来的理解障碍。
记录开放手术过程对教学和医疗评估至关重要,但传统单摄像头方法常因术者头身遮挡及固定角度限制,导致视频内容不清晰。本研究采用六视角多摄像头系统,通过提出一种全监督时序预测方法,从多个同步录制视频流中选取最佳画面序列,确保每个时刻的最优视角。模型利用预训练模型提取手术视频的视觉与语义特征,经由包含TimeBlocks的时序网络捕捉序列依赖关系,再通过线性嵌入层降维,最终由Softmax分类器根据最高概率选出最优摄像机视图。实验构建了五组同步六视角开放甲状腺切除术视频数据集,结果表明该方法在长时预测下仍保持较高准确率,且优于当前主流时序预测技术。本文首次提出该创新框架,显著提升手术视频分析能力,对改进手术教学与患者安全具有重要意义。
原文摘要 · Abstract (English)
Recording the open surgery process is essential for educational and medical evaluation purposes; however, traditional single-camera methods often face challenges such as occlusions caused by the surgeon's head and body, as well as limitations due to fixed camera angles, which reduce comprehensibility of the video content. This study addresses these limitations by employing a multi-viewpoint camera recording system, capturing the surgical procedure from six different angles to mitigate occlusions. We propose a fully supervised learning-based time series prediction method to choose the best shot sequences from multiple simultaneously recorded video streams, ensuring optimal viewpoints at each moment. Our time series prediction model forecasts future camera selections by extracting and fusing visual and semantic features from surgical videos using pre-trained models. These features are processed by a temporal prediction network with TimeBlocks to capture sequential dependencies. A linear embedding layer reduces dimensionality, and a Softmax classifier selects the optimal camera view based on the highest probability. In our experiments, we created five groups of open thyroidectomy videos, each with simultaneous recordings from six different angles. The results demonstrate that our method achieves competitive accuracy compared to traditional supervised methods, even when predicting over longer time horizons. Furthermore, our approach outperforms state-of-the-art time series prediction techniques on our dataset. This manuscript makes a unique contribution by presenting an innovative framework that advances surgical video analysis techniques, with significant implications for improving surgical education and patient safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。