用临床决策路径生成4063个医学问答数据,评估大模型推理能力
HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways
- 通过临床决策路径自动生成真实患者病例和问答对
- 覆盖17个医疗领域,每条数据含完整推理链条,共4063个案例
- 适合评估大模型在医疗场景下的多步推理与可解释性
HealthBranches 是一个新型医学问答基准数据集,专为评估大语言模型(LLMs)的复杂推理能力而设计。该数据集通过半自动化流程,将医学领域的明确决策路径转化为包含真实患者案例、问题与答案的结构化数据。涵盖17个医疗主题的4,063个病例研究,每个数据点均基于临床验证的推理链构建。支持开放问答与多选题形式,并首次完整包含每组问答的推理路径。其结构化设计可有效评估大模型在多步推理及检索增强生成(RAG)环境下的表现。该数据集为构建更可信、可解释、临床可靠的高风险领域大模型奠定基础,同时具备教学应用价值。
原文摘要 · Abstract (English)
HealthBranches is a novel benchmark dataset for medical Question-Answering (Q&A), specifically designed to evaluate complex reasoning in Large Language Models (LLMs). This dataset is generated through a semi-automated pipeline that transforms explicit decision pathways from medical source into realistic patient cases with associated questions and answers. Covering 4,063 case studies across 17 healthcare topics, each data point is based on clinically validated reasoning chains. HealthBranches supports both open-ended and multiple-choice question formats and uniquely includes the full reasoning path for each Q&A. Its structured design enables robust evaluation of LLMs' multi-step inference capabilities, including their performance in structured Retrieval-Augmented Generation (RAG) contexts. HealthBranches establishes a foundation for the development of more trustworthy, interpretable, and clinically reliable LLMs in high-stakes domains while also serving as a valuable resource for educational purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。