用AI自动识别手术反馈质量,提升教学效果评估效率
A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

- 采用多智能体提示+领域知识注入,挖掘可解释的反馈评分标准
- 在4200条真实反馈上验证,预测学员行为改变和导师认可度更准
- 适合医学教育研究者与临床教学评估系统开发者
主刀医生在术中提供的口头反馈对住院医师技能成长至关重要,但评估反馈质量及其对学员行为影响仍具挑战。以往研究依赖专家人工标注,建立宽泛分类,忽视清晰度、紧迫性等表达质量维度;现有自动化方法如关键词分析和主题建模也难以捕捉这些细微特征。本文提出一种两阶段大模型框架,通过多智能体提示与外科领域知识注入,发现一组人类可理解的反馈质量标准(如:鼓励性、紧迫性、清晰性)。基于这些标准,采用大模型作为裁判实现对实时手术反馈的自动评分。在4200个训练反馈实例上的评估表明,AI发现的标准在预测反馈有效性方面优于以往内容基框架,包括观察到的学员行为调整和导师认可度。本工作推动了手术室沟通质量的可扩展、以人为本的评估,为改进外科教学实践提供基础。
原文摘要 · Abstract (English)
Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, assessing the quality of trainer feedback and its effectiveness in influencing trainee behavior during live surgery remains a challenge. Prior studies assessed feedback content relying on extensive manual annotation by expert human raters and focused on developing broad taxonomies that overlook the qualitative aspects of feedback delivery such as clarity or urgency. Limited existing automated methods, including keyword analysis and topic modeling, also fail to capture these nuanced aspects. We introduce a two-stage LLM-based framework that discovers interpretable feedback quality criteria grounded in the context of surgical training. Our method uses multi-agent prompting and surgical domain knowledge injection to discover a small set of human interpretable scoring criteria (e.g., Encouraging, Urgent, Clear). These criteria are then used to automatically score live surgical feedback via an LLM-as-a-judge approach. Evaluation on 4.2k trainer feedback instances demonstrates that our AI-discovered criteria outperform prior content-based frameworks in predicting feedback effectiveness, including observed trainee behavioral adjustments and trainer approval. This work advances scalable, human-aligned assessment of communication quality in the operating room and provides a foundation for improving surgical teaching practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。