用社交媒体数据识别阿片类药物滥用的临床与社会影响,提升公共卫生监测能力。
Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- 构建基于第一人称叙述的标注数据集,聚焦临床与社会后果识别。
- 微调DeBERTa-large模型在宽松词级F1达0.61,优于大语言模型。
- 小样本下仍可实现有效识别,适合资源有限的医疗场景应用。
非医疗用途的阿片类药物使用是紧迫的公共健康挑战,其临床与社会后果常在传统医疗环境中被低估。社交媒体平台中用户坦率分享的第一人称经历,为揭示这些影响提供了宝贵但未充分利用的数据源。本研究提出命名实体识别(NER)框架,从相关社交叙事中提取两类自述后果:临床影响(如戒断、抑郁)和社会影响(如失业)。为此,我们引入了高质量的RedditImpacts 2.0数据集,具有优化的标注指南和对第一人称披露的重点关注,解决了先前工作的关键局限。我们在零样本与少样本上下文学习设置下评估了微调的编码器模型与最先进的大语言模型(LLMs)。微调后的DeBERTa-large模型在宽松词级F1上达到0.61(95%置信区间:0.43–0.62),在精确率、跨度准确率及任务规范遵循性方面持续优于LLMs。此外,我们证明在显著更少标注数据条件下即可实现强性能,强调了在资源受限环境下部署稳健模型的可行性。研究结果凸显领域特定微调对临床NLP任务的价值,并推动负责任的AI工具发展,以增强成瘾监测、提升可解释性并支持真实世界医疗决策。然而,最佳模型仍显著低于专家间一致性(Cohen's kappa: 0.81),表明在需要深度领域知识的任务中,专家智能与当前最先进的NER/AI能力之间仍存在明显差距。
原文摘要 · Abstract (English)
Nonmedical opioid use is an urgent public health challenge, with far-reaching clinical and social consequences that are often underreported in traditional healthcare settings. Social media platforms, where individuals candidly share first-person experiences, offer a valuable yet underutilized source of insight into these impacts. In this study, we present a named entity recognition (NER) framework to extract two categories of self-reported consequences from social media narratives related to opioid use: ClinicalImpacts (e.g., withdrawal, depression) and SocialImpacts (e.g., job loss). To support this task, we introduce RedditImpacts 2.0, a high-quality dataset with refined annotation guidelines and a focus on first-person disclosures, addressing key limitations of prior work. We evaluate both fine-tuned encoder-based models and state-of-the-art large language models (LLMs) under zero- and few-shot in-context learning settings. Our fine-tuned DeBERTa-large model achieves a relaxed token-level F1 of 0.61 [95% CI: 0.43-0.62], consistently outperforming LLMs in precision, span accuracy, and adherence to task-specific guidelines. Furthermore, we show that strong NER performance can be achieved with substantially less labeled data, emphasizing the feasibility of deploying robust models in resource-limited settings. Our findings underscore the value of domain-specific fine-tuning for clinical NLP tasks and contribute to the responsible development of AI tools that may enhance addiction surveillance, improve interpretability, and support real-world healthcare decision-making. The best performing model, however, still significantly underperforms compared to inter-expert agreement (Cohen's kappa: 0.81), demonstrating that a gap persists between expert intelligence and current state-of-the-art NER/AI capabilities for tasks requiring deep domain knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。