专用医疗记录AI比主流大模型更准,尤其在临床文档生成上表现突出。
Ambient AI Scribing Support: Comparing the Performance of Specialized AI Agentic Architecture to Leading Foundational Models
- 采用多智能体架构的专用AI模型,针对医疗记录优化设计。
- 召回率73.3%、精确率78.6%,综合得分高于GPT-4o等主流模型。
- 临床医生满意度高,适合医疗场景下的自动化文书工作。
本研究对比了专用于医疗记录的Sporo Health AI Scribe与多种大语言模型(GPT-4o、GPT-3.5、Gemma-9B、Llama-3.2-3B)在临床文档生成中的表现。基于合作诊所的匿名患者转录文本,以医师提供的SOAP笔记为基准,各模型通过零样本提示生成摘要,评估指标包括召回率、精确率和F1分数。Sporo在所有指标中均表现最佳,召回率73.3%、精确率78.6%、F1分数75.3%,且性能波动最小。统计检验显示其显著优于GPT-3.5、Gemma-9B和Llama-3.2-3B(p < 0.05),虽优于GPT-4o达10%,但差异未达显著水平(p = 0.25)。临床用户满意度调查(修正版PDQI-9)也表明Sporo输出更准确、相关性更高。结果凸显其多智能体架构在提升临床工作流中的潜力。
原文摘要 · Abstract (English)
This study compares Sporo Health's AI Scribe, a proprietary model fine-tuned for medical scribing, with various LLMs (GPT-4o, GPT-3.5, Gemma-9B, and Llama-3.2-3B) in clinical documentation. We analyzed de-identified patient transcripts from partner clinics, using clinician-provided SOAP notes as the ground truth. Each model generated SOAP summaries using zero-shot prompting, with performance assessed via recall, precision, and F1 scores. Sporo outperformed all models, achieving the highest recall (73.3%), precision (78.6%), and F1 score (75.3%) with the lowest performance variance. Statistically significant differences (p < 0.05) were found between Sporo and the other models, with post-hoc tests showing significant improvements over GPT-3.5, Gemma-9B, and Llama 3.2-3B. While Sporo outperformed GPT-4o by up to 10%, the difference was not statistically significant (p = 0.25). Clinical user satisfaction, measured with a modified PDQI-9 inventory, favored Sporo. Evaluations indicated Sporo's outputs were more accurate and relevant. This highlights the potential of Sporo's multi-agentic architecture to improve clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。