arXiv:2506.13363cs.CL2025-06被引 1

用强化学习仅用100样本实现高效医疗文档信息提取

Efficient Medical VIE via Reinforcement Learning

  • 基于可验证奖励的强化学习框架,仅需100标注样本
  • 在医学视觉信息抽取任务上显著提升准确率与召回率
  • 适合需要低资源、高精度医疗文本结构化的场景

视觉信息抽取(VIE)将非结构化文档图像转换为如JSON等结构化格式,在报告分析和在线问诊等医疗应用中至关重要。传统方法依赖OCR与语言模型,而端到端多模态模型可直接生成JSON。然而,领域特定的模式和高昂标注成本限制了其在医疗VIE中的效果。本文基于可验证奖励的强化学习(RLVR)框架,仅使用100个标注样本,通过增强数据多样性、平衡精确率-召回率奖励机制以减少幻觉并提升字段覆盖度,以及创新采样策略提升推理能力。对Qwen2.5-VL-7B进行微调后,在医疗VIE任务上达到当前最优性能,显著提升F1、精确率与召回率。尽管模型在相似任务上表现优异,但在不相似任务上性能下降,凸显领域适配的重要性。案例研究进一步验证了训练与推理中推理能力的价值。

原文摘要 · Abstract (English)

Visual Information Extraction (VIE) converts unstructured document images into structured formats like JSON, critical for medical applications such as report analysis and online consultations. Traditional methods rely on OCR and language models, while end-to-end multimodal models offer direct JSON generation. However, domain-specific schemas and high annotation costs limit their effectiveness in medical VIE. We base our approach on the Reinforcement Learning with Verifiable Rewards (RLVR) framework to address these challenges using only 100 annotated samples. Our approach ensures dataset diversity, a balanced precision-recall reward mechanism to reduce hallucinations and improve field coverage, and innovative sampling strategies to enhance reasoning capabilities. Fine-tuning Qwen2.5-VL-7B with our RLVR method, we achieve state-of-the-art performance on medical VIE tasks, significantly improving F1, precision, and recall. While our models excel on tasks similar to medical datasets, performance drops on dissimilar tasks, highlighting the need for domain-specific optimization. Case studies further demonstrate the value of reasoning during training and inference for VIE.

医疗AI视觉信息抽取强化学习低资源学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。