用视觉语言模型让癌症转诊单自动审核更可信、可追溯。
RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing

- 用多模态模型直接理解转诊单图文内容,避免传统文字识别漏洞。
- 微调后的Qwen3-VL模型准确率达96.1%,证据定位正确率60.6%。
- 适合医疗AI审计、临床决策支持系统开发者参考。
紧急疑似结直肠癌(CRC)转诊单因半结构化病历需人工审阅而造成流程瓶颈。原版RAPTOR系统依赖大语言模型进行结构化提取,但使用独立的OCR阶段,易受手写体、版式变化及视觉证据关联丢失影响。本文提出RAPTOR+,一种基于视觉语言模型(VLMs)的端到端转诊理解框架。我们在223份经临床标注的紧急转诊单上评估了微调的VLM、商用与开源零样本VLM,以及原始OCR流水线。引入一种关注视觉定位的评估框架,同时衡量提取准确率和证据定位能力。结果显示零样本模型存在明显定位缺口:Gemini 2.5 Flash阅读准确率达92.6%,但严格安全得分仅1.2%。相比之下,微调后的Qwen3-VL-8B达到96.1%阅读准确率与60.6%严格安全得分,显著提升可验证的视觉证据关联性。结果表明任务定制化微调对可靠、可审计的临床文档理解至关重要。RAPTOR+使提取的转诊决策可关联原始视觉证据,支持更安全高效的癌症转诊分诊。
原文摘要 · Abstract (English)
Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual review and transcription. The original RAPTOR system used Large Language Models for structured extraction but relied on a separate OCR stage, making it vulnerable to handwriting, layout variation, and loss of visual evidence linkage. We present RAPTOR+, a multimodal extension that uses Vision-Language Models (VLMs) for end-to-end referral understanding. We evaluate fine-tuned VLMs, commercial and open-source zero-shot VLMs, and the original OCR-based pipeline on 223 clinically curated CRC urgent referral forms. We also introduce a grounding-aware evaluation framework that measures both extraction accuracy and evidence localisation. Results show a clear grounding gap in zero-shot models. Gemini 2.5 Flash achieved 92.6% Reading Accuracy but only 1.2% Strict Safety. In contrast, fine-tuned Qwen3-VL-8B achieved 96.1% Reading Accuracy and 60.6% Strict Safety, substantially improving verifiable evidence grounding. These findings show that task-specific fine-tuning is essential for reliable, auditable clinical document understanding. RAPTOR+ enables extracted referral decisions to be linked to visual evidence, supporting safer and more efficient cancer referral triage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。