arXiv:2509.10744cs.CLcs.AI2025-09中稿 · publication at the…被引 2

用论文自动生成千道科学问答题,小模型靠推理痕迹也能超GPT-4。

Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models

  • 从2.2万篇论文自动构建1.6万道多选题,全流程可扩展。
  • 用GPT-4的推理轨迹做检索,小模型在癌症生物题上超越GPT-4。
  • 适合做小模型领域适应和科学知识评测的研究者。

随着科学知识以前所未有的速度增长,评估基准必须随之更新,以反映最新发现并确保语言模型在当前、多样化的文献上接受测试。我们提出一种可扩展、模块化的框架,直接从大规模科学论文语料库中生成多项选择题问答(MCQA)基准。该流水线自动化了MCQA生成的每个阶段,包括PDF解析、语义分块、问题生成和模型评估。作为案例研究,我们从2.2万篇开放获取的辐射与癌症生物学论文中生成了超过1.6万道多选题。随后,我们在这些题目上评估了一系列小型语言模型(1.1B-14B参数),比较基线准确率与从论文衍生的语义片段及从GPT-4.1提炼出的推理轨迹进行检索增强生成(RAG)的表现。结果表明,推理轨迹检索持续提升性能,在合成与专家标注基准上均实现显著进步,使多个小型模型在2023年天体辐射与癌症生物学考试中超越GPT-4。

原文摘要 · Abstract (English)

As scientific knowledge grows at an unprecedented pace, evaluation benchmarks must evolve to reflect new discoveries and ensure language models are tested on current, diverse literature. We propose a scalable, modular framework for generating multiple-choice question-answering (MCQA) benchmarks directly from large corpora of scientific papers. Our pipeline automates every stage of MCQA creation, including PDF parsing, semantic chunking, question generation, and model evaluation. As a case study, we generate more than 16,000 MCQs from 22,000 open-access articles in radiation and cancer biology. We then evaluate a suite of small language models (1.1B-14B parameters) on these questions, comparing baseline accuracy with retrieval-augmented generation (RAG) from paper-derived semantic chunks and from reasoning traces distilled from GPT-4.1. We find that reasoning-trace retrieval consistently improves performance on both synthetic and expert-annotated benchmarks, enabling several small models to surpass GPT-4 on the 2023 Astro Radiation and Cancer Biology exam.

科学问答小模型RAG自动构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。