arXiv:2603.11407cs.IR2026-03被引 1

用合成病历训练模型,精准提取癫痫发作频率信息。

Reproducible Synthetic Clinical Letters for Seizure Frequency Information Extraction

  • 生成结构化合成病历,模拟真实临床文本并标注发作频率
  • 仅用合成数据训练的模型在真实病历上达到0.847的精确率
  • 输出带证据支撑的结果,适合医疗研究与临床验证

癫痫发作频率对癫痫研究和临床诊疗至关重要,但通常记录在难以标注和共享的自由文本病历中。我们构建了一个可复现、保护隐私的框架,利用完全合成但任务忠实的癫痫病历提取发作频率信息。定义了涵盖常见发作负担描述的结构化标签体系,包括明确频率、范围、集群、无发作期、未知频率及明确无发作声明。使用教师语言模型生成类NHS风格的合成病历,附带标准化标签、推理过程和证据片段。在这些合成数据上微调多个开源大模型(参数量4B-14B),比较直接数值预测与结构化标签预测的效果,并测试基于证据的输出。在经医生校验的独立真实病历集上,仅使用合成数据训练的模型表现良好,结构化标签始终优于直接回归。使用15,000份合成训练数据,模型在细粒度类别上达到最高0.788的micro-F1,在实用类别上达0.847;医学导向的4B模型分别取得0.787和0.858。基于证据的输出支持快速临床核查与错误分析。结果表明,合成、结构化、基于证据的监督可实现无需共享敏感文本的稳健发作频率提取,并可能推广至其他时间复杂临床信息抽取任务。

原文摘要 · Abstract (English)

Seizure-frequency information is important for epilepsy research and clinical care, but it is usually recorded in variable free-text clinic letters that are hard to annotate and share. We developed a reproducible, privacy-preserving framework for extracting seizure frequency using fully synthetic yet task-faithful epilepsy letters. We defined a structured label scheme covering common descriptions of seizure burden, including explicit rates, ranges, clusters, seizure-free intervals, unknown frequency, and explicit no-seizure statements. A teacher language model generated NHS-style synthetic letters paired with normalized labels, rationales, and evidence spans. We fine-tuned several open-weight language models (4B-14B parameters) on these synthetic letters to extract seizure frequency from full documents, comparing direct numeric prediction with structured label prediction and testing evidence-grounded outputs. On a clinician-checked held-out set of real clinic letters, models trained only on synthetic data generalized well, and structured labels consistently outperformed direct numeric regression. With 15,000 synthetic training letters, models achieved micro-F1 scores up to 0.788 for fine-grained categories and 0.847 for pragmatic categories; a medically oriented 4B model achieved 0.787 and 0.858, respectively. Evidence-grounded outputs also supported rapid clinical verification and error analysis. These results show that synthetic, structured, evidence-grounded supervision can enable robust seizure-frequency extraction without sharing sensitive patient text and may generalize to other temporally complex clinical information extraction tasks.

临床信息抽取合成数据癫痫研究大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。