用小模型phi-3-mini自动识别医疗与运动伤文本,效果接近人工。
Evaluation of the phi-3-mini SLM for identification of texts related to medicine, health, and sports injuries
- 用phi-3-mini对百万级新闻标题做主题相关性评分并筛选样本。
- 在高过滤条件下,医疗文本与人工评分相关性达0.385(p<0.001)。
- 模型适合低资源场景下医疗类文本的初步筛选,不适用于高精度体育伤文本判断。
小型语言模型(SLMs)有望用于从文档和网络中自动标注与医学/健康相关的文本。由于资源需求远低于大语言模型(LLMs),SLMs可在更多设备上部署。尽管现有基准测试如MedQA常用于评估其性能,但实际应用表现可能因模型参数量而异。本研究将微软phi-3-mini-4k-instruct的文本主题相关性得分,与7名人类评估者对1144个医学/健康相关文本及1117个运动伤相关文本的评分进行对比。这些文本来自约900万条新闻标题,经由phi-3-mini-4k-instruct处理并打分。样本按低(1个布尔条件)或高(多个布尔条件)过滤标准选取。结果显示:低过滤下运动伤文本相关性为0.3413(p<0.001),高过滤下医疗文本为0.3854(p<0.001),低过滤下医疗文本为0.2255(p<0.001);高过滤下运动伤文本相关性仅为0.0318(p=0.4466),无显著关联。
原文摘要 · Abstract (English)
Small Language Models (SLMs) have potential to be used for automatically labelling and identifying aspects of text data for medicine/health-related purposes from documents and the web. As their resource requirements are significantly lower than Large Language Models (LLMs), these can be deployed potentially on more types of devices. SLMs often are benchmarked on health/medicine-related tasks, such as MedQA, although performance on these can vary especially depending on the size of the model in terms of number of parameters. Furthermore, these test results may not necessarily reflect real-world performance regarding the automatic labelling or identification of texts in documents and the web. As a result, we compared topic-relatedness scores from Microsofts phi-3-mini-4k-instruct SLM to the topic-relatedness scores from 7 human evaluators on 1144 samples of medical/health-related texts and 1117 samples of sports injury-related texts. These texts were from a larger dataset of about 9 million news headlines, each of which were processed and assigned scores by phi-3-mini-4k-instruct. Our sample was selected (filtered) based on 1 (low filtering) or more (high filtering) Boolean conditions on the phi-3 SLM scores. We found low-moderate significant correlations between the scores from the SLM and human evaluators for sports injury texts with low filtering (\r{ho} = 0.3413, p < 0.001) and medicine/health texts with high filtering (\r{ho} = 0.3854, p < 0.001), and low significant correlation for medicine/health texts with low filtering (\r{ho} = 0.2255, p < 0.001). There was negligible, insignificant correlation for sports injury-related texts with high filtering (\r{ho} = 0.0318, p = 0.4466).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。