构建医学论文偏倚风险评估基准,助力可靠文献问答。
Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark
- 基于系统综述的偏倚风险框架,构建三类评估任务。
- 涵盖500+医学研究,含专家对方法学的细粒度标注。
- 验证大模型在偏倚评估中推理与检索能力的关键作用。
能够通过审查科学文献回答问题的系统正变得日益可行。为得出可靠结论,这些系统应考虑不同研究证据的质量,对采用有效方法的研究给予更高权重。本文提出一个衡量生物医学论文方法学强度的基准,借鉴系统综述中使用的偏倚风险框架。该基准源于500余项生物医学研究,包含三类任务,涵盖专家对研究方法学的判断,包括对研究内部偏倚风险的评估。基准包含经人工验证的标注流程,实现评审意见与论文句子的细粒度对齐。分析表明,大语言模型的推理与检索能力显著影响其在偏倚评估中的表现。数据集已公开于 https://github.com/RoBBR-Benchmark/RoBBR。
原文摘要 · Abstract (English)
Systems that answer questions by reviewing the scientific literature are becoming increasingly feasible. To draw reliable conclusions, these systems should take into account the quality of available evidence from different studies, placing more weight on studies that use a valid methodology. We present a benchmark for measuring the methodological strength of biomedical papers, drawing on the risk-of-bias framework used for systematic reviews. Derived from over 500 biomedical studies, the three benchmark tasks encompass expert reviewers' judgments of studies' research methodologies, including the assessments of risk of bias within these studies. The benchmark contains a human-validated annotation pipeline for fine-grained alignment of reviewers' judgments with research paper sentences. Our analyses show that large language models' reasoning and retrieval capabilities impact their effectiveness with risk-of-bias assessment. The dataset is available at https://github.com/RoBBR-Benchmark/RoBBR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。