用词性标注和高效微调提升社交媒体健康提及识别准确率
Enhancing Health Mention Classification Performance: A Study on Advancements in Parameter Efficient Tuning
- 结合词性标注与参数高效微调技术
- 在三个数据集上F1得分优于当前最优方法
- 小模型实现高效训练,适合资源受限场景
健康提及分类(HMC)在利用社交媒体实时追踪公共卫生事件中至关重要。然而,由于健康提及常含隐喻语言和描述性术语,而非直接表达个人疾病,分类难度大。本文提出通过增强生物医学自然语言处理模型的参数,结合词性标注(POS)信息与参数高效微调(PEFT)技术,在三个主流数据集(RHDM、PHM、Illness)上进行实验。结果表明,引入POS信息并使用PEFT显著提升F1分数,且在更小模型和更低训练成本下表现优于现有方法。研究证实,该方法能有效提升社交媒体中健康提及的分类精度,同时优化模型规模与训练效率。
原文摘要 · Abstract (English)
Health Mention Classification (HMC) plays a critical role in leveraging social media posts for real-time tracking and public health monitoring. Nevertheless, the process of HMC presents significant challenges due to its intricate nature, primarily stemming from the contextual aspects of health mentions, such as figurative language and descriptive terminology, rather than explicitly reflecting a personal ailment. To address this problem, we argue that clearer mentions can be achieved through conventional fine-tuning with enhanced parameters of biomedical natural language methods (NLP). In this study, we explore different techniques such as the utilisation of part-of-speech (POS) tagger information, improving on PEFT techniques, and different combinations thereof. Extensive experiments are conducted on three widely used datasets: RHDM, PHM, and Illness. The results incorporated POS tagger information, and leveraging PEFT techniques significantly improves performance in terms of F1-score compared to state-of-the-art methods across all three datasets by utilising smaller models and efficient training. Furthermore, the findings highlight the effectiveness of incorporating POS tagger information and leveraging PEFT techniques for HMC. In conclusion, the proposed methodology presents a potentially effective approach to accurately classifying health mentions in social media posts while optimising the model size and training efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。