通过分析网络问药行为,提前预警潜在健康危机。
When Curiosity Signals Danger: Predicting Health Crises Through Online Medication Inquiries
- 构建了标注药物问题严重性的新数据集。
- 传统与大模型方法均能有效识别高危提问。
- 适合医疗安全、AI辅助诊疗研究者参考。
在线医疗论坛是患者用药关切的丰富来源,其中部分问题可能暗示困惑、误用,甚至早期健康危机信号。本研究构建了一个新型标注数据集,从论坛中提取药物相关问题,并根据临床风险因素人工标注其严重性。我们对比了六种传统机器学习分类器(基于TF-IDF)和三种基于大语言模型(LLM)的先进分类方法,评估其在识别高危问题上的表现。结果表明,经典与现代方法均具备支持数字健康领域实时分诊与预警系统的潜力。所构建的数据集已公开,以促进患者生成数据、自然语言处理与重大健康事件早期预警研究的融合。数据集与基准测试链接:https://github.com/Dvora-coder/LLM-Medication-QA-Risk-Classifier-MediGuard。
原文摘要 · Abstract (English)
Online medical forums are a rich and underutilized source of insight into patient concerns, especially regarding medication use. Some of the many questions users pose may signal confusion, misuse, or even the early warning signs of a developing health crisis. Detecting these critical questions that may precede severe adverse events or life-threatening complications is vital for timely intervention and improving patient safety. This study introduces a novel annotated dataset of medication-related questions extracted from online forums. Each entry is manually labelled for criticality based on clinical risk factors. We benchmark the performance of six traditional machine learning classifiers using TF-IDF textual representations, alongside three state-of-the-art large language model (LLM)-based classification approaches that leverage deep contextual understanding. Our results highlight the potential of classical and modern methods to support real-time triage and alert systems in digital health spaces. The curated dataset is made publicly available to encourage further research at the intersection of patient-generated data, natural language processing, and early warning systems for critical health events. The dataset and benchmark are available at: https://github.com/Dvora-coder/LLM-Medication-QA-Risk-Classifier-MediGuard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。