用大模型从社交媒体提取药物副作用,构建知识图谱分析司美格鲁肽安全风险。
Crowdsourcing-Based Knowledge Graph Construction for Drug Side Effects Using Large Language Models with an Application on Semaglutide
- 用大语言模型从Reddit提取未结构化用药反馈,自动构建副作用知识图谱。
- 发现不同品牌司美格鲁肽的副作用报告随时间变化趋势,部分与FAERS数据库一致。
- 为医患提供真实世界用药体验补充信息,适合药物安全研究者参考。
社交媒体是药学警戒中获取患者真实体验的重要数据源,但其非结构化和噪声多的特点使数据挖掘困难。本文提出一种系统性框架,利用大语言模型(LLMs)从社交媒体中提取药物副作用,并组织成知识图谱(KG)。以用于减肥的司美格鲁肽为例,基于Reddit数据构建知识图谱,开展跨品牌、跨时间的副作用报告分析。结果通过与美国食品药品监督管理局不良事件报告系统(FAERS)数据对比验证,揭示了患者中心视角下的安全性特征,补充了现有知识体系。本研究证明了使用大模型将社交媒体数据转化为结构化知识图谱在药学警戒中的可行性。
原文摘要 · Abstract (English)
Social media is a rich source of real-world data that captures valuable patient experience information for pharmacovigilance. However, mining data from unstructured and noisy social media content remains a challenging task. We present a systematic framework that leverages large language models (LLMs) to extract medication side effects from social media and organize them into a knowledge graph (KG). We apply this framework to semaglutide for weight loss using data from Reddit. Using the constructed knowledge graph, we perform comprehensive analyses to investigate reported side effects across different semaglutide brands over time. These findings are further validated through comparison with adverse events reported in the FAERS database, providing important patient-centered insights into semaglutide's side effects that complement its safety profile and current knowledge base of semaglutide for both healthcare professionals and patients. Our work demonstrates the feasibility of using LLMs to transform social media data into structured KGs for pharmacovigilance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。