arXiv:2412.00760eess.AScs.AI2024-12NeurIPS被引 2

自动分析手术训练中的口语反馈,提升教学评估效率。

Automating Feedback Analysis in Surgical Training: Detection, Categorization, and Assessment

  • 融合语音检测与说话人分离,重建真实手术对话
  • 误检率降低14%,反馈识别F1达0.79
  • 可支持行为调整预测与技术反馈分类,媲美人工标注

本文提出首个从真实手术录音中重构手术对话的框架,对刻画教学任务至关重要。在手术培训中,导师在术中提供的口头反馈对保障安全、即时纠正行为及长期技能习得极为关键,但其非结构化与专业性强的特点使分析量化困难。自动化系统对大规模处理此类复杂信息不可或缺,可构建结构化数据集以改进反馈分析和手术教育。本框架整合语音活动检测、说话人分离与自动语音识别,并创新性地实现两点:1)消除因手术室噪声导致的语音识别幻觉(不存在的语句);2)仅用少量语音样本区分导师与学员发言。基于33场真实手术数据验证,系统能有效重建教学对话并检测反馈实例(F1=0.79±0.07),其中幻觉消除使反馈检测性能提升约14%。在下游临床相关任务上——预测学员行为调整与分类技术反馈——表现分别达到F1=0.82±0.03与0.81±0.03,与人工标注相当。结果表明该框架在支持临床任务方面高效且优于传统人工方法。

原文摘要 · Abstract (English)

This work introduces the first framework for reconstructing surgical dialogue from unstructured real-world recordings, which is crucial for characterizing teaching tasks. In surgical training, the formative verbal feedback that trainers provide to trainees during live surgeries is crucial for ensuring safety, correcting behavior immediately, and facilitating long-term skill acquisition. However, analyzing and quantifying this feedback is challenging due to its unstructured and specialized nature. Automated systems are essential to manage these complexities at scale, allowing for the creation of structured datasets that enhance feedback analysis and improve surgical education. Our framework integrates voice activity detection, speaker diarization, and automated speech recaognition, with a novel enhancement that 1) removes hallucinations (non-existent utterances generated during speech recognition fueled by noise in the operating room) and 2) separates speech from trainers and trainees using few-shot voice samples. These aspects are vital for reconstructing accurate surgical dialogues and understanding the roles of operating room participants. Using data from 33 real-world surgeries, we demonstrated the system's capability to reconstruct surgical teaching dialogues and detect feedback instances effectively (F1 score of 0.79+/-0.07). Moreover, our hallucination removal step improves feedback detection performance by ~14%. Evaluation on downstream clinically relevant tasks of predicting Behavioral Adjustment of trainees and classifying Technical feedback, showed performances comparable to manual annotations with F1 scores of 0.82+/0.03 and 0.81+/0.03 respectively. These results highlight the effectiveness of our framework in supporting clinically relevant tasks and improving over manual methods.

手术训练语音分析反馈检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。