用标签筛选案例提升LLM生成医疗事故原因与预防措施的准确性和稳定性。
Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

- 基于数据集标签选择少样本示例,增强LLM对医疗事故分析的指导性。
- 在JMID数据集上,该方法使生成结果精确度最高且安全过滤触发最少。
- 适合医疗AI系统开发、临床决策支持场景,尤其关注生成可靠性时使用。
在医疗等高风险领域,大型语言模型(LLMs)的可靠性至关重要,尤其是在从事故报告中生成临床洞察时。本研究提出一种基于标签的少样本示例选择方法,用于引导LLM从医疗事故细节中生成背景/因果因素及预防措施。实验采用日本医疗事故数据集(JMID),包含3,884条真实世界医疗事故与未遂事件报告,其标注涵盖广泛标签(如“药物”、“输血治疗”)。对比三种少样本示例选择策略——随机采样、余弦相似度选择与本文提出的标签方法,在GPT-4o和LLaMA 3.3上进行测试。结果表明,标签方法在生成精度上表现最佳且行为最稳定,而基于相似度的选择常引发非预期输出并触发安全过滤。研究提示,利用人类可理解的数据集标签选择示例,可显著提升临床LLM应用中的生成精度与稳定性。
原文摘要 · Abstract (English)
In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights from incident reports. This study proposes a tag-based few-shot example selection method for prompting LLMs to generate background/causal factors and preventive measures from details of the medical incidents. For our experiments, we use the Japanese Medical Incident Dataset (JMID), a structured dataset of 3,884 real-world medical accident and near-miss reports. These reports are variably annotated with a wide range of tags--some include descriptive information (e.g., "medications," "blood transfusion therapy"). We compare three few-shot example selection strategies--random sampling, cosine similarity-based selection, and our proposed tag-based method--using GPT-4o and LLaMA 3.3. Results show that the tag-based approach achieves the highest precision and most stable generation behavior, while similarity-based selection often leads to unintended outputs and safety filter activation. These findings suggest that selecting examples based on human-interpretable dataset tags can improve generation precision and stability in clinical LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。