arXiv:2608.23061cs.AI2026-08

用混合模型提升超声报告O-RADS分级准确率,接近专家水平。

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

  • 分步处理:先提取特征再按指南判断,避免直接推理出错。
  • 最高准确率达99.2%(387/390),加权kappa值1.00,接近完美一致。
  • 适合临床医生和医学AI研发者,可减少误判并提升报告可信度。

背景:利用大语言模型(LLMs)自动化基于临床指南的决策仍面临可靠性差、幻觉多和可解释性弱等挑战。本研究比较了多种LLM及推理策略在自动从自由文本盆腔超声报告中进行卵巢-附件报告与数据系统(O-RADS)分类的表现。方法:回顾性纳入连续接受盆腔超声检查的卵巢肿块患者。测试了8种LLM,采用三种推理策略:隐含知识端到端、规则引导端到端,以及将特征提取与规则分类解耦的基于特征的混合架构。参考标准为专家共识确立的O-RADS分类。结果:共评估310名女性,390个卵巢肿块。基于特征的混合架构结合Gemini 3.6 Flash表现最佳,准确率达99.2%(387/390),与参考标准几乎完全一致(加权kappa=1.00;95%置信区间:0.99–1.00)。其性能优于原始临床报告(准确率87.7% [342/390];加权kappa=0.94;95%CI: 0.91–0.96)及端到端LLM策略(准确率范围65.6% [256/390] 至 95.9% [374/390])。在结构化特征提取方面,Gemini 3.6 Flash整体准确率高于Claude Fable 5(98.9% vs 97.8%;P<0.001)。混合架构减少了误分类错误,并缓解了原始报告中过度分阶的问题。结论:将临床特征提取与确定性指南执行分离的基于特征的混合型LLM架构,可实现高精度、高可靠性和高可解释性的自动化O-RADS分类,为标准化指南驱动的临床决策提供有前景的新路径。

原文摘要 · Abstract (English)

Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for automated Ovarian-Adnexal Reporting and Data System (O-RADS) classification from free-text pelvic ultrasound reports. Methods: In this retrospective study, consecutive patients with ovarian masses who underwent pelvic ultrasound were included. Eight LLMs were tested with three reasoning strategies: implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture that decoupled feature extraction from rule-based classification. The reference standard was O-RADS categorization established by expert consensus. Results: A total of 310 women with 390 ovarian masses were evaluated. The feature-based hybrid architecture using Gemini 3.6 Flash demonstrated the best performance, achieving an accuracy of 99.2% (387 of 390) and almost perfect agreement with the reference standard (weighted kappa = 1.00; 95% CI: 0.99-1.00). Its performance surpassed that of original clinical reports (accuracy, 87.7% [342 of 390]; weighted kappa = 0.94; 95% CI: 0.91-0.96) and end-to-end LLM strategies (accuracy range, 65.6% [256 of 390] to 95.9% [374 of 390]). For structured feature extraction, Gemini 3.6 Flash demonstrated higher overall accuracy than Claude Fable 5 (98.9% vs 97.8%; P < 0.001). The hybrid architecture reduced misclassification errors and mitigated the overstaging tendency observed in original reports. Conclusion: The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.

医学AIO-RADS大模型分类系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。