arXiv:2511.15159cs.CVcs.AI2025-11中稿 · ML4H 2025

用手术动作结构生成真实医生风格的反馈,提升准确性与可信度。

Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation

  • 从真实手术反馈中提取仪器-动作-目标三元组,构建标准化动作语义体系
  • 通过视频与时间序列运动信息提升动作识别,使准确率显著提高
  • 用结构化表示引导大模型生成符合临床标准的反馈,更易被医生验证

高质量术中反馈对医学生技能提升至关重要。自动化生成类人反馈可实现大规模、及时、一致的指导,但需模型理解临床相关表征。本文提出一种结构感知流程:从33例真实导师-学员对话中学习手术动作本体,并用于条件化反馈生成。贡献包括:(1)从真实反馈文本中挖掘仪器-动作-目标(IAT)三元组,将表面形式聚类为标准化类别;(2)微调视频到IAT模型,融合手术流程、任务上下文及细粒度时间运动信息;(3)展示如何有效利用IAT三元组指导GPT-4o生成临床可信、类人风格的反馈。实验表明:在视频到IAT识别任务中,引入上下文和时序追踪使AUC提升(仪器:0.67→0.74;动作:0.60→0.63;组织:0.74→0.79)。在反馈生成任务中(1-5分评分制,1=错误/危险,3=可接受,5=完美匹配),仅基于视频的GPT-4o得分为2.17,而加入IAT条件后达2.44(+12.4%),可接受生成比例从21%升至42%。传统文本相似性指标也提升:词错误率下降15%-31%,ROUGE值上升9%-64%。在显式IAT结构基础上生成反馈,显著提高内容真实性,支持临床可验证与审计使用。

原文摘要 · Abstract (English)

High-quality intraoperative feedback from a surgical trainer is pivotal for improving trainee performance and long-term skill acquisition. Automating natural, trainer-style feedback promises timely, accessible, and consistent guidance at scale but requires models that understand clinically relevant representations. We present a structure-aware pipeline that learns a surgical action ontology from real trainer-to-trainee transcripts (33 surgeries) and uses it to condition feedback generation. We contribute by (1) mining Instrument-Action-Target (IAT) triplets from real-world feedback text and clustering surface forms into normalized categories, (2) fine-tuning a video-to-IAT model that leverages the surgical procedure and task contexts as well as fine-grained temporal instrument motion, and (3) demonstrating how to effectively use IAT triplet representations to guide GPT-4o in generating clinically grounded, trainer-style feedback. We show that, on Task 1: Video-to-IAT recognition, our context injection and temporal tracking deliver consistent AUC gains (Instrument: 0.67 to 0.74; Action: 0.60 to 0.63; Tissue: 0.74 to 0.79). For Task 2: feedback text generation (rated on a 1-5 fidelity rubric where 1 = opposite/unsafe, 3 = admissible, and 5 = perfect match to a human trainer), GPT-4o from video alone scores 2.17, while IAT conditioning reaches 2.44 (+12.4%), doubling the share of admissible generations with score >= 3 from 21% to 42%. Traditional text-similarity metrics also improve: word error rate decreases by 15-31% and ROUGE (phrase/substring overlap) increases by 9-64%. Grounding generation in explicit IAT structure improves fidelity and yields clinician-verifiable rationales, supporting auditable use in surgical training.

手术生成自然语言医疗AI反馈系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。