arXiv:2605.04221cs.CLcs.AI2026-05

小模型自动生成提示词,精准提取牙科病历隐私信息

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction

  • 小模型自主生成、验证并优化特定实体的提示词
  • 微平均和宏平均F1分别达0.864和0.837(经偏好优化后)
  • 适合需要本地部署、保护隐私的医疗信息提取场景

从牙科病历中进行临床命名实体识别面临文档高度非结构化、领域特定性强及隐私敏感等挑战。我们开发了一个可本地部署的框架,使小型语言模型能够自动生成、验证、优化和评估用于提取多种临床实体的专用提示词。基于1,200条标注病历,我们评估了多个开源权重模型,并采用多提示集成推理,进一步通过QLoRA微调与直接偏好优化(DPO)适配选定模型。模型性能差异显著,凸显任务特定评估的重要性,而非依赖通用基准。Qwen2.5-14B-Instruct表现最佳,经DPO后,其微平均和宏平均F1得分分别达到0.864和0.837;Llama-3.1-8B-Instruct则分别为0.806和0.797。结果表明,自动提示优化结合轻量级偏好训练,可支持在本地部署的小模型实现可扩展的临床信息提取。

原文摘要 · Abstract (English)

Clinical named entity recognition from dental progress notes is challenging because documentation is highly unstructured, domain-specific, and often privacy-sensitive. We developed a locally deployable framework that enables small language models to self-generate, verify, refine, and evaluate entity-specific prompts for extracting multiple clinical entities from dental notes. Using 1,200 annotated notes, we evaluated candidate open-weight models with multi-prompt ensemble inference and further adapted selected models using QLoRA-based supervised fine-tuning and direct preference optimization. Model performance varied substantially, highlighting the need for task-specific evaluation rather than reliance on generic benchmarks. Qwen2.5-14B-Instruct achieved the strongest baseline performance. After DPO, Qwen2.5-14B-Instruct and Llama-3.1-8B-Instruct achieved micro/macro F1 scores of 0.864/0.837 and 0.806/0.797, respectively. These findings suggest that automated prompt optimization combined with lightweight preference-based post-training can support scalable clinical information extraction using locally deployed small language models.

小模型提示工程医疗信息提取隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。