LLMs在精神类药物不良反应应对上难达专家水平
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use
- 构建Psych-ADR基准与ADRA评估框架
- 仅70.86%策略与专家一致,行动建议少12.32%
- 虽语气契合但内容复杂难懂,适合高风险场景评估
精神类药物的不良反应(ADRs)是心理健康患者住院的主要原因。当前医疗系统和在线社区在处理此类问题时存在局限,大型语言模型(LLMs)或可填补这一空白。尽管LLMs能力不断提升,但其在识别精神类药物相关不良反应及提供有效缓解策略方面的能力尚未被系统研究。为此,我们提出Psych-ADR基准和不良药物反应响应评估(ADRA)框架,用于系统评估LLMs在检测不良反应表达和生成专家对齐缓解策略方面的表现。分析显示,LLMs难以理解不良反应的细微差别,也难以区分不同类型。虽然在情感和语气上与专家一致,但其回应更复杂、更难理解,且仅有70.86%的策略与专家一致。此外,平均而言,其建议的可操作性低12.32%。本工作为高风险领域中以策略为导向的任务提供了全面的评估基准与框架。
原文摘要 · Abstract (English)
Adverse Drug Reactions (ADRs) from psychiatric medications are the leading cause of hospitalizations among mental health patients. With healthcare systems and online communities facing limitations in resolving ADR-related issues, Large Language Models (LLMs) have the potential to fill this gap. Despite the increasing capabilities of LLMs, past research has not explored their capabilities in detecting ADRs related to psychiatric medications or in providing effective harm reduction strategies. To address this, we introduce the Psych-ADR benchmark and the Adverse Drug Reaction Response Assessment (ADRA) framework to systematically evaluate LLM performance in detecting ADR expressions and delivering expert-aligned mitigation strategies. Our analyses show that LLMs struggle with understanding the nuances of ADRs and differentiating between types of ADRs. While LLMs align with experts in terms of expressed emotions and tone of the text, their responses are more complex, harder to read, and only 70.86% aligned with expert strategies. Furthermore, they provide less actionable advice by a margin of 12.32% on average. Our work provides a comprehensive benchmark and evaluation framework for assessing LLMs in strategy-driven tasks within high-risk domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。