arXiv:2503.16464cs.HCcs.AI2025-03

用AI聊天平台辅助多学科医生讨论复杂心脏病病例,大幅缩短总结时间。

Human-Centered AI in Multidisciplinary Medical Discussions: Evaluating the Feasibility of a Chat-Based Approach to Case Assessment

  • 构建医患协作的AI聊天系统,模拟真实多学科病例讨论。
  • AI总结使信息提炼时间平均减少79.98%,整体幻觉率仅3.62%。
  • 适合临床决策支持、医疗AI落地及多学科会诊场景使用。

本研究探讨了基于聊天平台的人类中心人工智能(AI)在多学科医学讨论中的可行性,聚焦于患有多种慢性病的心血管疾病患者。我们基于既往病例报告、医疗差错报告及医生实际经验构建了五个模拟病例,通过与医师合作评估平台的可行性、效率提升及讨论内容量化。分析显示,使用AI进行总结可使信息提炼时间平均减少79.98%。整体幻觉率在1.01%至5.73%之间,平均为3.62%;有害幻觉率在0.00%至2.09%之间,平均为0.49%。形态学分析表明,多学科评估比单医生评估能更复杂、更详细地呈现医学知识。通过知识图谱的中心性指标比较,发现多学科评估具有更高的结构复杂性。结果表明,AI辅助总结显著降低讨论时间,同时保持知识结构完整性,支持以人类为中心的聊天式多学科决策模式的可行性。

原文摘要 · Abstract (English)

In this study, we investigate the feasibility of using a human-centered artificial intelligence (AI) chat platform where medical specialists collaboratively assess complex cases. As the target population for this platform, we focus on patients with cardiovascular diseases who are in a state of multimorbidity, that is, suffering from multiple chronic conditions. We evaluate simulated cases with multiple diseases using a chat application by collaborating with physicians to assess feasibility, efficiency gains through AI utilization, and the quantification of discussion content. We constructed simulated cases based on past case reports, medical errors reports and complex cases of cardiovascular diseases experienced by the physicians. The analysis of discussions across five simulated cases demonstrated a significant reduction in the time required for summarization using AI, with an average reduction of 79.98\%. Additionally, we examined hallucination rates in AI-generated summaries used in multidisciplinary medical discussions. The overall hallucination rate ranged from 1.01\% to 5.73\%, with an average of 3.62\%, whereas the harmful hallucination rate varied from 0.00\% to 2.09\%, with an average of 0.49\%. Furthermore, morphological analysis demonstrated that multidisciplinary assessments enabled a more complex and detailed representation of medical knowledge compared with single physician assessments. We examined structural differences between multidisciplinary and single physician assessments using centrality metrics derived from the knowledge graph. In this study, we demonstrated that AI-assisted summarization significantly reduced the time required for medical discussions while maintaining structured knowledge representation. These findings can support the feasibility of AI-assisted chat-based discussions as a human-centered approach to multidisciplinary medical decision-making.

AI医疗多学科会诊知识图谱对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。