用大模型辅助自闭症临床诊断,效果接近专业医生。
Copiloting Diagnosis of Autism in Real Clinical Scenarios via LLMs
- 构建基于大模型的自闭症诊断助手,结合评分与解释。
- 诊断准确率达F1=81.79%(二分类),最小误差仅0.4643。
- 揭示大模型在精神健康应用中的优势与局限,指导未来落地。
自闭症谱系障碍(ASD)是一种严重影响个体日常生活与社交参与的广泛性发育障碍。尽管已有大量研究致力于支持ASD的临床诊断,但基于大语言模型(LLMs)的方法在真实临床场景下的系统性探索仍显不足,尤其针对自闭症诊断观察量表第二版(ADOS-2)的应用。为此,我们提出名为ADOS-Copilot的框架,在评分与解释之间取得平衡,并探究影响LLMs在此任务中表现的关键因素。实验结果表明,该框架在诊断性能上可与临床医生相媲美:最小平均绝对误差(MAE)为0.4643,二分类F1得分为81.79%,三分类F1得分为78.37%。此外,我们从ADOS-2特性、大模型能力、语言表达及模型规模等角度,系统剖析了当前大模型在该任务中的优势与局限,旨在推动其在更广泛精神健康领域的应用。期望更多研究能走向真实临床实践,为特殊儿童打开一扇温暖之窗。
原文摘要 · Abstract (English)
Autism spectrum disorder(ASD) is a pervasive developmental disorder that significantly impacts the daily functioning and social participation of individuals. Despite the abundance of research focused on supporting the clinical diagnosis of ASD, there is still a lack of systematic and comprehensive exploration in the field of methods based on Large Language Models (LLMs), particularly regarding the real-world clinical diagnostic scenarios based on Autism Diagnostic Observation Schedule, Second Edition (ADOS-2). Therefore, we have proposed a framework called ADOS-Copilot, which strikes a balance between scoring and explanation and explored the factors that influence the performance of LLMs in this task. The experimental results indicate that our proposed framework is competitive with the diagnostic results of clinicians, with a minimum MAE of 0.4643, binary classification F1-score of 81.79\%, and ternary classification F1-score of 78.37\%. Furthermore, we have systematically elucidated the strengths and limitations of current LLMs in this task from the perspectives of ADOS-2, LLMs' capabilities, language, and model scale aiming to inspire and guide the future application of LLMs in a broader fields of mental health disorders. We hope for more research to be transferred into real clinical practice, opening a window of kindness to the world for eccentric children.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。