arXiv:2605.08045cs.CL2026-05中稿 · ISBI 2026被引 1

将心脏磁共振报告转为结构化数据并给出置信度,提升医疗数据可用性。

Uncertainty-Aware Structured Data Extraction from Full CMR Reports via Distilled LLMs

  • 用教师-学生蒸馏技术实现离线推理,减少人工标注。
  • 字段级准确率达99.65%,且输出可信度评分。
  • 适合临床研究、电子病历系统和质量控制场景。

将自由文本的心脏磁共振(CMR)报告转化为可审计的结构化数据,仍是队列构建、纵向数据维护和临床决策支持的瓶颈。我们提出CMR-EXTR,一个轻量级框架,可将自由文本的CMR报告转换为结构化数据,并为每个字段生成置信度分数以支持质量控制。通过教师-学生蒸馏流程,实现完全离线推理,同时降低人工标注需求。不确定性评估融合三个互补原则——分布合理性、采样稳定性与跨字段一致性——用于优先安排人工审核。实验表明,CMR-EXTR在变量级别达到99.65%的准确率,既保证了可靠的抽取效果,也提供了有信息量的置信度评分。据我们所知,这是首个集成置信度估计的专用于CMR报告的提取系统。代码已开源:https://github.com/yuyi1005/CMR-EXTR。

原文摘要 · Abstract (English)

Converting free-text cardiac magnetic resonance (CMR) reports into auditable structured data remains a bottleneck for cohort assembly, longitudinal curation, and clinical decision support. We present CMR-EXTR, a lightweight framework that converts free-text CMR reports into structured data and assigns per-field confidence for quality control. A teacher-student distillation pipeline enables fully offline inference while limiting manual annotation. Uncertainty integrates three complementary principles -- distribution plausibility, sampling stability, and cross-field consistency -- to triage human review. Experiments show that CMR-EXTR achieves 99.65% variable-level accuracy, demonstrating both reliable extraction and informative confidence scores. To our knowledge, this is the first CMR-specific extraction system with integrated confidence estimation. The code is available at https://github.com/yuyi1005/CMR-EXTR.

医疗文本结构化提取置信度自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。