arXiv:2607.05880cs.CVcs.AI2026-07

HR1.5可基于影像、病史和先验信息自动生成放射科报告,减轻医生负担。

Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context

论文配图:Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
图 1 · 摘自论文原文
  • 多模态大模型融合图像与文本输入,生成结构化报告
  • 在模拟FRCR考试中唯一达标,多项指标领先
  • 支持可解释性分析,适合临床落地评估

影像需求增长速度远超放射科人力扩充能力,仅靠培训和招聘无法解决报告积压问题。最直接的突破口是减少放射科医生撰写报告的时间与精力——这项工作需解读图像、整合病史与既往研究,并撰写结构化发现。我们提出霍里森·辐射1.5(Harrison.Rad 1.5,HR1.5),一个专用于放射领域的多模态大语言模型,可接收交错的文本与视觉输入,生成涵盖平片放射学(包括计算机放射摄影、胸部、骨肌、腹部、脊柱、盆腔X线及乳腺钼靶)的结构化与非结构化报告。HR1.5通过三阶段训练流程:在放射科报告上对基础语言模型进行领域适配;在约600万图像-报告对上采用课程学习的困难负样本进行对比视觉编码器训练;以及在多轮对话上的视觉问答微调。我们采用扩展了本体同义词匹配与极性矛盾检测的Findings-Diagnosis评分框架,在RadBench(模拟FRCR 2B短病例考试,基于Angoff法设定通过阈值)、ReXGradient及内部多模态数据集上进行评估。HR1.5是唯一达到模拟FRCR通过标准的系统,在闭式临床问题准确率、各解剖区域、内部多部位及乳腺报告任务,以及公开胸部报告的主要临床相关得分上均表现最优。我们进一步分析了可解释性与模型行为,包括敏感问题的Grad-CAM热图、注意力分析与置信度估计,为未来临床应用提供负责任的评估框架。

原文摘要 · Abstract (English)

Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone. The most direct opportunity is reducing the time and effort radiologists spend producing reports, a task that requires interpreting images, integrating clinical history and prior studies, and drafting structured findings. We present Harrison.Rad 1.5 (HR1.5), a radiology-specific multimodal large language model that accepts interleaved text and visual inputs and generates structured and unstructured text across plain-film radiology, spanning computed radiography, chest, musculoskeletal, abdominal, spine, and pelvic x-rays, and mammography. HR1.5 is trained through a three-stage pipeline: domain adaptation of a base language model on radiology reports, contrastive vision-encoder training with curriculum-based hard negatives on ~6 million image-report instances, and visual-question-answering fine-tuning on multi-turn conversations. We evaluate it with a Findings-Diagnosis scoring framework that extends RadGraph-XL entity extraction with ontology-based synonym matching and polarity-contradiction detection, benchmarked on RadBench, a simulated FRCR 2B Short Case examination scored against Angoff-method thresholds, ReXGradient, and internal multi-modality datasets. HR1.5 is the only system evaluated to meet the simulated FRCR passing standard and achieves the highest accuracy on closed-format clinical questions, across anatomical regions, on internal multi-body-part and mammography reporting, and on the primary clinically-aligned score for public chest reporting. We further examine explainability and model behaviour, including question-sensitive Grad-CAM heatmaps, attention analysis, and confidence estimation, to support responsible future evaluation toward clinical use, and a framework for clinically grounded assessment of report quality.

医学AI报告生成多模态放射科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。