arXiv:2502.01789cs.AIcs.MA2025-02被引 8

用多智能体AI自动识别临床记录中的认知问题,效果媲美专家。

An Agentic AI Workflow for Detecting Cognitive Concerns in Real-world Data

  • 设计多智能体协作系统,自动解析3338份临床笔记中的认知线索。
  • F1分数达0.90,特异性高达1.00,比基准更快优化提示。
  • 适合需要高效筛查认知障碍的医疗场景,尤其资源有限地区。

早期识别认知问题至关重要,但常因症状微妙而被忽视。本研究开发并验证了一种全自动化多智能体AI工作流,基于LLaMA 3 8B模型,在来自麻省总医院布里格姆的3,338份临床笔记上识别认知问题。该智能体工作流通过任务特定智能体动态协作,从临床笔记中提取关键信息,并与专家驱动基准进行比较。两者均表现出高分类性能,F1分数分别为0.90和0.91。多智能体工作流展现出更高的特异性(1.00),并在更少迭代次数内完成提示优化。尽管在验证数据上性能略有下降,该工作流仍保持完美特异性。结果表明,全自动多智能体AI工作流可实现专家级准确率,且效率更高,为临床环境中检测认知问题提供了可扩展、低成本的解决方案。

原文摘要 · Abstract (English)

Early identification of cognitive concerns is critical but often hindered by subtle symptom presentation. This study developed and validated a fully automated, multi-agent AI workflow using LLaMA 3 8B to identify cognitive concerns in 3,338 clinical notes from Mass General Brigham. The agentic workflow, leveraging task-specific agents that dynamically collaborate to extract meaningful insights from clinical notes, was compared to an expert-driven benchmark. Both workflows achieved high classification performance, with F1-scores of 0.90 and 0.91, respectively. The agentic workflow demonstrated improved specificity (1.00) and achieved prompt refinement in fewer iterations. Although both workflows showed reduced performance on validation data, the agentic workflow maintained perfect specificity. These findings highlight the potential of fully automated multi-agent AI workflows to achieve expert-level accuracy with greater efficiency, offering a scalable and cost-effective solution for detecting cognitive concerns in clinical settings.

多智能体认知筛查临床AI自动化诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。