arXiv:2505.06912cs.CV2025-05被引 5

构建3万+临床问答对,让AI推理过程可验证可信。

Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI

  • 用人类+LLM协作流程生成带专家验证的推理链
  • 产出31,247个跨领域的医学问答对,含专家审核的思考过程
  • 适合想提升医疗AI可解释性与可信度的研究者

尽管大型语言模型在医学问答中表现强劲,但其透明的‘黑箱’推理过程严重阻碍了临床应用,限制了医生信任。当前医疗LLM多依赖科学文献或合成数据,缺乏细致的专家验证和高临床相关性,难以真正提升专业能力。为此,我们提出一个高度临床相关的数据集,包含31,247个医学问答对,每个都配有专家验证的思维链(CoT)解释。该数据集通过可扩展的人类-LLM混合流程构建:由LLM生成推理,再经医疗专家依据结构化标准逐轮评审、评分并修正,不达标内容通过人工重写或引导式LLM再生,直至达成专家共识。该公开数据集为发展具备透明、可验证推理能力的医疗LLM提供了关键资源,推动更安全、可解释的医学AI发展。

原文摘要 · Abstract (English)

Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is compounded by the predominant reliance of current medical LLMs on corpora from scientific literature or synthetic data, which often lack the granular expert validation and high clinical relevance essential for advancing their specialized medical capabilities. To address these critical gaps, we introduce a highly clinically relevant dataset with 31,247 medical question-answer pairs, each accompanied by expert-validated chain-of-thought (CoT) explanations. This resource, spanning multiple clinical domains, was curated via a scalable human-LLM hybrid pipeline: LLM-generated rationales were iteratively reviewed, scored, and refined by medical experts against a structured rubric, with substandard outputs revised through human effort or guided LLM regeneration until expert consensus. This publicly available dataset provides a vital source for the development of medical LLMs that capable of transparent and verifiable reasoning, thereby advancing safer and more interpretable AI in medicine.

医疗AI可解释性数据集思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。