让专家高效参与提示词优化,提升医疗文本分类效果
Keeping Experts in the Loop: Expert-Guided Optimization for Clinical Data Classification using Large Language Models
- 设计迭代采样算法,精准识别需专家反馈的关键案例
- 在无需大量标注的情况下,显著提升分类F1分数
- 适合希望降低专家负担又保证模型性能的医疗AI团队
自大型语言模型(LLMs)出现以来,如何有效利用其在医疗领域的潜力成为核心挑战。从非结构化临床笔记中提取洞察的关键障碍在于提示词工程。尽管其对任务性能至关重要,但缺乏明确的提示优化框架。现有方法要么依赖人工精细调整,耗时且难扩展;要么完全自动化,未能充分吸纳领域专家价值。为此,我们提出StructEase框架,通过SamplEase迭代采样算法,识别出专家反馈能带来显著性能提升的高价值案例,从而最小化专家介入,有效改善分类结果。该方法减少重复标注、降低人为错误,提升分类效果。我们在美国国家电子伤害监测系统(NEISS)的去标识临床叙事数据集上评估了StructEase,相比现有方法实现显著性能提升。研究证明,将专家融入LLM工作流具有重要价值,在保持极低专家投入的同时显著提高F1分数。StructEase凭借透明性、灵活性与可扩展性,为医疗及其他领域中融合专家知识的LLM应用奠定了基础。
原文摘要 · Abstract (English)
Since the emergence of Large Language Models (LLMs), the challenge of effectively leveraging their potential in healthcare has taken center stage. A critical barrier to using LLMs for extracting insights from unstructured clinical notes lies in the prompt engineering process. Despite its pivotal role in determining task performance, a clear framework for prompt optimization remains absent. Current methods to address this gap take either a manual prompt refinement approach, where domain experts collaborate with prompt engineers to create an optimal prompt, which is time-intensive and difficult to scale, or through employing automatic prompt optimizing approaches, where the value of the input of domain experts is not fully realized. To address this, we propose StructEase, a novel framework that bridges the gap between automation and the input of human expertise in prompt engineering. A core innovation of the framework is SamplEase, an iterative sampling algorithm that identifies high-value cases where expert feedback drives significant performance improvements. This approach minimizes expert intervention, to effectively enhance classification outcomes. This targeted approach reduces labeling redundancy, mitigates human error, and enhances classification outcomes. We evaluated the performance of StructEase using a dataset of de-identified clinical narratives from the US National Electronic Injury Surveillance System (NEISS), demonstrating significant gains in classification performance compared to current methods. Our findings underscore the value of expert integration in LLM workflows, achieving notable improvements in F1 score while maintaining minimal expert effort. By combining transparency, flexibility, and scalability, StructEase sets the foundation for a framework to integrate expert input into LLM workflows in healthcare and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。