用AI检测医学教材中的不当用语,提升教学内容的包容性与安全性。
AI-Powered Detection of Inappropriate Language in Medical School Curricula
- 用微调的小模型和提示工程的大模型识别医疗教材中的不当语言。
- 多标签分类模型在标注数据上表现最佳,加入未标记文本后AUC提升25%。
- 适合医学教育工作者、课程审查者及AI伦理研究者参考使用。
不当语言(如过时、排他性或非以患者为中心的术语)在医学教学材料中的使用会显著影响临床训练、患者互动及健康结果。尽管这些材料历史悠久,但许多仍包含当前医学标准下被认为不恰当的内容。由于课程内容量大,手动识别不当语言及其子类别成本高昂且不现实。为此,我们首次评估了微调的小语言模型(SLMs)和采用上下文学习的预训练大模型(LLMs)在约500份文档、超过12,000页数据上的表现。针对SLMs,我们测试了通用分类器、子类别的二分类器、多标签分类器以及两阶段分层检测流程;针对LLMs,我们尝试了包含子类定义和/或示例的提示变体。结果显示,即使经过精心设计的示例,LLama-3 8B和70B均显著低于小模型表现。多标签分类器在标注数据上表现最优,而通过引入未标记文本作为负样本进行训练,可使特定分类器的AUC最高提升25%,成为最有效的模型用于消除医学课程中的有害语言。
原文摘要 · Abstract (English)
The use of inappropriate language -- such as outdated, exclusionary, or non-patient-centered terms -- medical instructional materials can significantly influence clinical training, patient interactions, and health outcomes. Despite their reputability, many materials developed over past decades contain examples now considered inappropriate by current medical standards. Given the volume of curricular content, manually identifying instances of inappropriate use of language (IUL) and its subcategories for systematic review is prohibitively costly and impractical. To address this challenge, we conduct a first-in-class evaluation of small language models (SLMs) fine-tuned on labeled data and pre-trained LLMs with in-context learning on a dataset containing approximately 500 documents and over 12,000 pages. For SLMs, we consider: (1) a general IUL classifier, (2) subcategory-specific binary classifiers, (3) a multilabel classifier, and (4) a two-stage hierarchical pipeline for general IUL detection followed by multilabel classification. For LLMs, we consider variations of prompts that include subcategory definitions and/or shots. We found that both LLama-3 8B and 70B, even with carefully curated shots, are largely outperformed by SLMs. While the multilabel classifier performs best on annotated data, supplementing training with unflagged excerpts as negative examples boosts the specific classifiers' AUC by up to 25%, making them most effective models for mitigating harmful language in medical curricula.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。