arXiv:2608.30021cs.LGcs.AI2026-08

对比大模型与专用模型,发现领域训练比模型规模更重要。

Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models

论文配图:Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models
图 1 · 摘自论文原文
  • 用领域数据训练小型BERT模型,精准检测癌症PET/CT报告错误
  • 1500万参数模型达94.4%准确率,远超最强提示大模型的84.0%
  • 专用模型更高效,适合临床部署,尤其适合资源有限场景

放射科报告中的错误可能影响患者治疗,但自动质量保障仍具挑战,因错误往往细微且需专业判断。尽管大语言模型(LLMs)已被用于报告验证,其在胸部X光之外的临床场景表现仍不明确。为此,我们首次系统评估了语言模型在PET/CT报告错误检测中的能力,比较了轻量级领域专用模型与SOTA开源大模型(Qwen3-32B、Gemma-3-27B、Llama-3.3-70B)。研究基于23名放射科医生十年间收集的30,633份肿瘤学FDG PET/CT报告,训练领域特定BERT模型以识别临床相关的合成错误,并在11,500份保留测试集上评估。一个1500万参数的模型实现94.4%的平衡准确率和5.8%的假阳性率,优于最强提示大模型的84.0%。对Llama-3.3-70B进行任务微调后性能提升至94.4%,但计算开销显著更高。结果表明,领域训练比模型规模更关键,支持采用紧凑模型实现高精度、低耗能的自动化报告质检。

原文摘要 · Abstract (English)

Errors in radiology reports can adversely affect patient treatment, yet automated report quality assurance remains challenging because errors are often subtle and require domain expertise to detect. Although large language models (LLMs) have recently been proposed for radiology report verification, their ability to detect clinically meaningful errors beyond chest X-ray datasets remains under-explored. To this end, we present the first systematic evaluation of language models for PET/CT report error detection, comparing compact domain-specific models with SOTA open-weight LLMs. We collected 30,633 oncology FDG PET/CT reports from 23 radiologists over 10 years. We trained domain-specific BERT models to detect clinically motivated synthetic reporting errors and evaluated alongside zero-/few-shot Qwen3-32B, Gemma-3-27B and Llama-3.3-70B on a held-out benchmark of 11,500 reports. A 15M-parameter model achieved 94.4% balanced accuracy with a 5.8% false-positive rate, compared with 84.0% for the strongest prompted LLM. Task-specific adaptation of Llama-3.3-70B closed this performance gap (94.4%) but retained substantially greater computational requirements. Our results suggest that domain-specific training matters more than model scale for PET/CT report error detection, supporting compact models as an accurate and computationally efficient approach to automated radiology report quality assurance.

医学影像错误检测大模型PET/CT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。