arXiv:2601.03232cs.CLcs.AI2026-01被引 2

构建多类医学影像报告标准数据集,评测41个大模型的报告标准化能力。

Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models

  • 用合成报告+医生双轮审核构建10类标准报告数据集
  • 32B级开源模型在提示引导下准确率达70%以上,接近闭源模型
  • 提示工程显著提升模型表现,复杂标准下性能下降明显

背景:报告与数据系统(RADS)用于标准化放射学风险沟通,但基于叙述性报告自动分配RADS标签面临指南复杂、输出格式限制及缺乏跨框架和模型规模的基准测试挑战。目的:创建经放射科医生验证的合成多RADS基准数据集RXL-RADSet,并比较41个量化的小型语言模型(SLMs)与一个专有模型在RADS分配上的有效性与准确性。方法:RXL-RADSet包含1,600份合成放射科报告,覆盖10种RADS(BI-RADS、CAD-RADS、GB-RADS、LI-RADS、Lung-RADS、NI-RADS、O-RADS、PI-RADS、TI-RADS、VI-RADS)及多种成像模态。报告由大语言模型根据情景计划和模拟放射科医生风格生成,并经过两阶段医生验证。评估了41个量化SLMs(12个系列,参数量0.135-32B)和GPT-5.2,在固定引导提示下的表现。主要终点为有效性和准确性;次要分析比较了引导提示与零样本提示的效果。结果:在引导提示下,GPT-5.2的有效性达99.8%,准确性为81.1%(1,600次预测)。聚合的SLMs(65,600次预测)有效性为96.8%,准确性为61.1%;20-32B范围内的顶级SLMs达到约99%有效性与中高70%准确性。性能随模型规模提升(在<1B与≥10B间出现拐点),并因RADS复杂性增加而下降,主要源于分类难度而非无效输出。引导提示较零样本提示显著提升有效性(99.2% vs 96.7%)与准确性(78.5% vs 69.6%)。结论:RXL-RADSet提供了经医生验证的多类RADS基准;大规模开源模型(20-32B)在引导提示下可接近专有模型性能,但在高复杂度方案上仍存差距。

原文摘要 · Abstract (English)

Background: Reporting and Data Systems (RADS) standardize radiology risk communication but automated RADS assignment from narrative reports is challenging because of guideline complexity, output-format constraints, and limited benchmarking across RADS frameworks and model sizes. Purpose: To create RXL-RADSet, a radiologist-verified synthetic multi-RADS benchmark, and compare validity and accuracy of open-weight small language models (SLMs) with a proprietary model for RADS assignment. Materials and Methods: RXL-RADSet contains 1,600 synthetic radiology reports across 10 RADS (BI-RADS, CAD-RADS, GB-RADS, LI-RADS, Lung-RADS, NI-RADS, O-RADS, PI-RADS, TI-RADS, VI-RADS) and multiple modalities. Reports were generated by LLMs using scenario plans and simulated radiologist styles and underwent two-stage radiologist verification. We evaluated 41 quantized SLMs (12 families, 0.135-32B parameters) and GPT-5.2 under a fixed guided prompt. Primary endpoints were validity and accuracy; a secondary analysis compared guided versus zero-shot prompting. Results: Under guided prompting GPT-5.2 achieved 99.8% validity and 81.1% accuracy (1,600 predictions). Pooled SLMs (65,600 predictions) achieved 96.8% validity and 61.1% accuracy; top SLMs in the 20-32B range reached ~99% validity and mid-to-high 70% accuracy. Performance scaled with model size (inflection between <1B and >=10B) and declined with RADS complexity primarily due to classification difficulty rather than invalid outputs. Guided prompting improved validity (99.2% vs 96.7%) and accuracy (78.5% vs 69.6%) compared with zero-shot. Conclusion: RXL-RADSet provides a radiologist-verified multi-RADS benchmark; large SLMs (20-32B) can approach proprietary-model performance under guided prompting, but gaps remain for higher-complexity schemes.

医学AI大模型评测多模态数据提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。