arXiv:2604.13077cs.CL2026-04

用大模型从葡萄牙语冠脉造影报告中提取生理指标,验证了可行性与局限性。

Can Large Language Models Reliably Extract Physiology Index Values from Coronary Angiography Reports?

论文配图:Can Large Language Models Reliably Extract Physiology Index Values from Coronary Angiography Reports?
图 1 · 摘自论文原文
  • 采用零样本提示和正则表达式后处理,实现非结构化报告的自动解析。
  • 在1342份报告上测试,最佳模型(Llama)准确率达87.6%,但约束生成降低性能。
  • 适合医疗数据挖掘、临床研究者关注,尤其对多语言医学文本处理有参考价值。

冠状动脉造影(CAG)报告包含重要的生理测量信息,但通常以非结构化自然语言形式存在,限制了其在研究中的应用。本研究探讨大型语言模型(LLMs)从葡萄牙语CAG报告中自动提取生理指标及其解剖位置的可行性。据我们所知,这是首个针对大规模(1342份)CAG报告语料库进行生理指标提取的研究,也是少数聚焦于CAG或葡萄牙语临床文本的工作。我们测试了本地隐私保护的通用及医学领域LLM,在不同提示策略下表现:零样本、少样本及包含不合理示例的少样本提示。同时引入约束生成与基于正则表达式的后处理步骤。由于测量值稀疏,提出多阶段评估框架,区分格式正确性、数值检测与数值准确性,并考虑临床误判成本不对称性。结果表明,非医学模型表现相当,最佳为Llama零样本提示(准确率87.6%),GPT-OSS对提示变化最鲁棒。MedGemma表现接近非医学模型,而MedLlama在无约束设置中结果格式错误,约束设置下性能显著下降。改变提示策略或添加正则表达式层未带来显著提升,但约束生成虽降低效果,却支持无法遵循模板的模型使用。

原文摘要 · Abstract (English)

Coronary angiography (CAG) reports contain clinically relevant physiological measurements, yet this information is typically in the form of unstructured natural language, limiting its use in research. We investigate the use of Large Language Models (LLMs) to automatically extract these values, along with their anatomical locations, from Portuguese CAG reports. To our knowledge, this study is the first addressing physiology indexes extraction from a large (1342 reports) corpus of CAG reports, and one of the few focusing on CAG or Portuguese clinical text. We explore local privacy-preserving general-purpose and medical LLMs under different settings. Prompting strategies included zero-shot, few-shot, and few-shot prompting with implausible examples. In addition, we apply constrained generation and introduce a post-processing step based on RegEx. Given the sparsity of measurements, we propose a multi-stage evaluation framework separating format validity, value detection, and value correctness, while accounting for asymmetric clinical error costs. This study demonstrates the potential of LLMs in for extracting physiological indices from Portuguese CAG reports. Non-medical models performed similarly, the best results were obtained with Llama with a zero-shot prompting, while GPT-OSS demonstrated the highest robustness to changes in the prompts. While MedGemma demonstrated similar results to non-medical models, MedLlama's results were out-of-format in the unconstrained setting, and had a significant lower performance in the constrained one. Changes in the prompt techinique and adding a RegEx layer showed no significant improvement across models, while using constrained generation decreased performance, although having the benefit of allowing the usage of specific models that are not able to conform with the templates.

医学文本大模型信息抽取葡萄牙语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。