扩充12000+道医考题,评估大模型医疗推理能力
HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
- 基于十年西班牙医考题构建多语言医疗推理数据集
- 模型规模与内在推理能力决定性能上限
- 适合研究医疗AI、大模型推理的开发者使用
我们推出HEAD-QA v2,是Vilares和Gómez-Rodríguez(2019)发布的西/英双语医疗多选题推理数据集的扩展与更新版本。该更新响应了对高质量、能捕捉医疗推理语言与概念复杂性的数据集日益增长的需求。我们扩展数据集至超过12,000道题目,覆盖十年西班牙专业医师考试内容,使用提示工程、RAG及基于概率的答案选择方法,对多个开源大模型进行基准测试,并提供额外的多语言版本以支持后续研究。结果表明,模型表现主要由模型规模和内在推理能力驱动,复杂推理策略带来的增益有限。这些结果共同确立HEAD-QA v2作为推动生物医学推理研究与模型改进的可靠资源。
原文摘要 · Abstract (English)
We introduce HEAD-QA v2, an expanded and updated version of a Spanish/English healthcare multiple-choice reasoning dataset originally released by Vilares and Gómez-Rodríguez (2019). The update responds to the growing need for high-quality datasets that capture the linguistic and conceptual complexity of healthcare reasoning. We extend the dataset to over 12,000 questions from ten years of Spanish professional exams, benchmark several open-source LLMs using prompting, RAG, and probability-based answer selection, and provide additional multilingual versions to support future work. Results indicate that performance is mainly driven by model scale and intrinsic reasoning ability, with complex inference strategies obtaining limited gains. Together, these results establish HEAD-QA v2 as a reliable resource for advancing research on biomedical reasoning and model improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。