用强化学习提升临床试验论文的量化推理能力,更准确判断研究结论。
Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning
- 基于数值推理框架,从论文中提取事件数等结构化数据
- 在CochraneForest上比检索系统高21%的F1分数
- 适合医疗证据合成、AI辅助决策等专业场景
医学系统评价在循证决策中至关重要,需整合多篇研究结果。自动化过程的核心瓶颈在于提取数值证据并确定特定结局和比较的研究级结论。以往方法将此问题视为文本蕴含任务,依赖表面文本线索,难以捕捉专家评估背后的量化逻辑。本文将其重构为定量推理问题:不依赖表层文本,而是提取结构化数值证据(如事件数、标准差),结合领域知识推导结论。我们构建了一个包含数值数据提取模型与效应量估算组件的推理系统,并采用监督微调和新设计的价值奖励模型进行强化学习训练。在CochraneForest基准测试中,使用强化学习训练的小型数值提取模型相比检索基线系统,F1分数提升达21个百分点;在RCTs基准上,优于参数超过4000亿的通用大模型最高9个百分点。结果表明,基于推理的方法在自动化系统评价合成中具有巨大潜力。
原文摘要 · Abstract (English)
Systematic reviews in medicine play a critical role in evidence-based decision-making by aggregating findings from multiple studies. A central bottleneck in automating this process is extracting numeric evidence and determining study-level conclusions for specific outcomes and comparisons. Prior work has framed this problem as a textual inference task by retrieving relevant content fragments and inferring conclusions from them. However, such approaches often rely on shallow textual cues and fail to capture the underlying numeric reasoning behind expert assessments. In this work, we conceptualise the problem as one of quantitative reasoning. Rather than inferring conclusions from surface text, we extract structured numerical evidence (e.g., event counts or standard deviations) and apply domain knowledge informed logic to derive outcome-specific conclusions. We develop a numeric reasoning system composed of a numeric data extraction model and an effect estimate component, enabling more accurate and interpretable inference aligned with the domain expert principles. We train the numeric data extraction model using different strategies, including supervised fine-tuning (SFT) and reinforcement learning (RL) with a new value reward model. When evaluated on the CochraneForest benchmark, our best-performing approach -- using RL to train a small-scale number extraction model -- yields up to a 21% absolute improvement in F1 score over retrieval-based systems and outperforms general-purpose LLMs of over 400B parameters by up to 9% on the RCTs benchmark. Our results demonstrate the promise of reasoning-driven approaches for automating systematic evidence synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。