用论文图文结构生成带证据的医学问答数据,提升模型准确性。
Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers

- 自动提取论文中的图文、表格与文本证据,构建完整支持链。
- 训练后在LAB-Bench上达48.0%准确率,比基线高12.6个百分点。
- 适合需要可解释医学AI的研究者和开发者使用。
通用视觉语言模型在生物医学研究中仍不可靠,因科学答案常分散于图表、表格、图注和正文引用段落中。现有后训练流程受限于昂贵的人工标注及丢失证据结构的合成数据。我们提出Ryze,一个全自动化系统,将原始生物医学论文转化为富含证据的训练数据集与领域专用的VLM。Ryze生成包含完整支持证据(视觉元素、图注、提取结构与引用段落)的问答对,通过图表感知提取和大模型清洗减少版面与OCR错误,并采用进度门控的后训练策略,结合监督微调与强化学习。基于Qwen3-VL-8B,Ryze以不足200美元成本训练出BioVLM-8B,在LAB-Bench上达到48.0%加权准确率,较基线提升12.6个百分点,超越GPT-5.2达3.8个百分点。我们开源Ryze及训练好的BioVLM-8B模型。
原文摘要 · Abstract (English)
General-purpose VLMs remain unreliable for biomedical research because valid answers in scientific papers depend on evidence split across figures, tables, charts, captions, and referring text. Existing post-training pipelines are bottlenecked by costly expert annotation and by synthetic data that drops this evidence structure. We present Ryze, a fully automated system that converts raw biomedical papers into an evidence-enriched training set and a domain-specialized VLM. Ryze synthesizes QA pairs with complete supporting evidence (visual element, caption, extracted structure, and referring paragraphs), reduces layout and OCR errors via chart/table-aware extraction and LLM-based cleansing, and applies a progress-gated post-training strategy combining supervised fine-tuning with reinforcement learning. Starting from Qwen3-VL-8B, Ryze produces BioVLM-8B at under USD 200, achieving 48.0% weighted accuracy on LAB-Bench, outperforming the base model by +12.6 percentage points (pp) and surpassing GPT-5.2 by +3.8 pp. We release Ryze as open source together with the trained BioVLM-8B model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。