arXiv:2603.10313cs.CL2026-03

大模型能精准识别社交媒体中的毒品俚语,提升毒瘾监测效果。

Large language models can disambiguate opioid slang on social media

  • 用大模型解析毒品俚语在上下文中的真实含义,解决歧义问题。
  • 在多个任务中,模型准确率远超传统词典方法(最高达0.972)。
  • 适合公共卫生、流行病学研究者用于低频话题的内容挖掘。

社交媒体文本在监测阿片类药物过量危机趋势方面具有潜力;然而,绝大多数社交媒体内容与阿片类药物无关。当前常用策略是使用阿片类药物相关术语词典作为筛选标准,但许多俚语如“smack”或“blues”具有常见非阿片类含义,造成歧义。大语言模型(LLMs)的高级文本推理能力为大规模消歧提供了可能。我们设计了三项任务评估四种先进模型(GPT-4、GPT-5、Gemini 2.5 Pro 和 Claude Sonnet 4.5):基于词典的任务(在给定帖子中判断特定术语是否指阿片类药物)、无词典任务(仅凭上下文识别阿片类相关内容)以及新兴俚语任务(识别模拟新俚语的阿片类相关内容)。所有模型在各项任务中表现优异。在基于词典任务中,模型F1得分('fenty'子任务:0.824–0.972;'smack'子任务:0.540–0.862)显著高于最佳词典策略(分别为0.126和0.009)。在无词典任务中,模型F1得分(0.544–0.769)也优于词典(0.080–0.540),且召回率更高。在新兴俚语任务中,所有模型的平均准确率(0.784)、F1(0.712)、精确率(0.981)和召回率(0.587)均超过所测试的两个词典。结果表明,大模型可用于识别低频主题内容,包括但不限于阿片类药物引用,从而增强下游分析与预测模型的数据质量。

原文摘要 · Abstract (English)

Social media text shows promise for monitoring trends in the opioid overdose crisis; however, the overwhelming majority of social media text is unrelated to opioids. When leveraging social media text to monitor trends in the ongoing opioid overdose crisis, a common strategy for identifying relevant content is to use a lexicon of opioid-related terms as inclusion criteria. However, many slang terms for opioids, such as "smack" or "blues," have common non-opioid meanings, making them ambiguous. The advanced textual reasoning capability of large language models (LLMs) presents an opportunity to disambiguate these slang terms at scale. We present three tasks on which to evaluate four state-of-the-art LLMs (GPT-4, GPT-5, Gemini 2.5 Pro, and Claude Sonnet 4.5): a lexicon-based setting, in which the LLM must disambiguate a specific term within the context of a given post; a lexicon-free setting, in which the LLM must identify opioid-related posts from context without a lexicon; and an emergent slang setting, in which the LLM must identify opioid-related posts with simulated new slang terms. All four LLMs showed excellent performance across all tasks. In both subtasks of the lexicon-based setting, LLM F1 scores ("fenty" subtask: 0.824-0.972; "smack" subtask: 0.540-0.862) far exceeded those of the best lexicon strategy (0.126 and 0.009, respectively). In the lexicon-free task, LLM F1 scores (0.544-0.769) surpassed those of lexicons (0.080-0.540), and LLMs demonstrated uniformly higher recall. On emergent slang, all LLMs had higher accuracy (average: 0.784), F1 score (average: 0.712), precision (average: 0.981), and recall (average: 0.587) than the two lexicons assessed. Our results show that LLMs can be used to identify relevant content for low-prevalence topics, including but not limited to opioid references, enhancing data provided to downstream analyses and predictive models.

大模型毒品监测自然语言处理社会媒体分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。