用视觉语言模型精准解析医疗文档,解决标注混乱难题
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
- 通过训练驱动的标签优化提升噪声标注质量
- 在医疗发票数据集上超越主流OCR与视觉语言模型
- 适合需要高精度字段提取的医疗信息化场景
医疗文档光学字符识别因复杂版式、专业术语和噪声标注而具有挑战性,且要求字段级精确匹配。现有OCR系统和通用视觉语言模型难以可靠解析此类文档。我们提出MeDocVL,一种面向查询驱动的医疗文档解析的后训练视觉语言模型。框架结合训练驱动的标签精炼方法,从噪声标注中构建高质量监督信号;并采用噪声感知的混合后训练策略,融合强化学习与监督微调,实现鲁棒且精确的信息提取。在医疗发票基准测试中,MeDocVL持续优于传统OCR系统和强基准视觉语言模型,在噪声监督下达到当前最优性能。
原文摘要 · Abstract (English)
Medical document OCR is challenging due to complex layouts, domain-specific terminology, and noisy annotations, while requiring strict field-level exact matching. Existing OCR systems and general-purpose vision-language models often fail to reliably parse such documents. We propose MeDocVL, a post-trained vision-language model for query-driven medical document parsing. Our framework combines Training-driven Label Refinement to construct high-quality supervision from noisy annotations, with a Noise-aware Hybrid Post-training strategy that integrates reinforcement learning and supervised fine-tuning to achieve robust and precise extraction. Experiments on medical invoice benchmarks show that MeDocVL consistently outperforms conventional OCR systems and strong VLM baselines, achieving state-of-the-art performance under noisy supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。