arXiv:2603.21494cs.CLcs.MA2026-03

用多智能体AI系统自动完成脑瘤随访评估,准确率超人工初判18.5个百分点。

Agentic Automation of BT-RADS Scoring: End-to-End Multi-Agent System for Standardized Brain Tumor Follow-up Assessment

  • 分角色智能体解析病历与影像,融合临床信息和肿瘤体积做决策
  • 整体准确率达76.0%,对关键类别如BT-1b、BT-4识别灵敏度极高
  • 适合放射科医生用于提升随访评估一致性,尤其擅长识别高危病例

脑瘤报告与数据系统(BT-RADS)标准化了弥漫性胶质瘤患者治疗后MRI应答评估,但需整合影像趋势、药物影响和放疗时间等复杂因素。本研究评估了一种端到端多智能体大语言模型(LLM)与卷积神经网络(CNN)结合的自动化系统。该系统在单中心连续509例治疗后胶质瘤MRI检查中进行回顾性验证。提取智能体从非结构化病历中识别临床变量(类固醇状态、贝伐单抗状态、放疗日期),评分智能体则结合提取变量与体积测量结果应用BT-RADS决策逻辑。独立认证的神经放射科专家作为参考标准。509例中492例符合纳入标准。系统分类准确率为374/492(76.0%;95% CI, 72.1%-79.6%),显著高于初始临床评估的283/492(57.5%;95% CI, 53.1%-61.8%),提升18.5个百分点(P<.001)。上下文依赖类别敏感度高(BT-1b 100%,BT-1a 92.7%,BT-3a 87.5%),阈值依赖类别中等(BT-3c 74.8%,BT-2 69.2%,BT-4 69.3%,BT-3b 57.1%)。BT-4阳性预测值达92.9%。多智能体系统相比初始临床评分,与专家标准一致性更高,对上下文依赖评分准确率高,且对BT-4检测具有高阳性预测值。

原文摘要 · Abstract (English)

The Brain Tumor Reporting and Data System (BT-RADS) standardizes post-treatment MRI response assessment in patients with diffuse gliomas but requires complex integration of imaging trends, medication effects, and radiation timing. This study evaluates an end-to-end multi-agent large language model (LLM) and convolutional neural network (CNN) system for automated BT-RADS classification. A multi-agent LLM system combined with automated CNN-based tumor segmentation was retrospectively evaluated on 509 consecutive post-treatment glioma MRI examinations from a single high-volume center. An extractor agent identified clinical variables (steroid status, bevacizumab status, radiation date) from unstructured clinical notes, while a scorer agent applied BT-RADS decision logic integrating extracted variables with volumetric measurements. Expert reference standard classifications were established by an independent board-certified neuroradiologist. Of 509 examinations, 492 met inclusion criteria. The system achieved 374/492 (76.0%; 95% CI, 72.1%-79.6%) accuracy versus 283/492 (57.5%; 95% CI, 53.1%-61.8%) for initial clinical assessments (+18.5 percentage points; P<.001). Context-dependent categories showed high sensitivity (BT-1b 100%, BT-1a 92.7%, BT-3a 87.5%), while threshold-dependent categories showed moderate sensitivity (BT-3c 74.8%, BT-2 69.2%, BT-4 69.3%, BT-3b 57.1%). For BT-4, positive predictive value was 92.9%. The multi-agent LLM system achieved higher BT-RADS classification agreement with expert reference standard compared to initial clinical scoring, with high accuracy for context-dependent scores and high positive predictive value for BT-4 detection.

脑瘤评估多智能体自动化诊断医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。