arXiv:2511.13381cs.CL2025-11

评测12个大模型在儿科临床中的表现,发现其诊断与人文关怀仍有短板。

Can Large Language Models Function as Qualified Pediatricians? A Systematic Evaluation in Real-World Clinical Contexts

  • 构建真实临床场景的评估框架PEDIASBench,覆盖19个儿科亚专科。
  • 顶尖模型在基础题上准确率超90%,复杂推理能力下降约15%。
  • 动态诊疗和伦理判断表现弱,适合辅助决策而非独立行医。

随着大语言模型在医学领域的快速发展,一个关键问题是它们能否在真实临床环境中胜任儿科医生角色。我们构建了PEDIASBench系统评估框架,基于知识体系,面向真实临床环境,从基础医学知识、动态诊疗能力、儿科医疗安全与伦理三个维度评估12个近两年发布的代表性模型,涵盖19个儿科亚专科和211种典型疾病。顶尖模型在基础题上表现优异,如Qwen3-235B-A22B在执照级问题中准确率超过90%,但任务复杂度上升后性能下降约15%,暴露出复杂推理能力不足;多选题测试显示整合推理与知识回忆存在缺陷。在动态诊疗场景中,DeepSeek-R1案例推理得分最高(均值0.58),但多数模型难以适应实时患者变化。伦理与安全任务中,Qwen2.5-72B表现最佳(准确率92.05%),但人文关怀仍显薄弱。结果表明,当前儿科大模型受限于动态决策能力和人文素养,未来需加强多模态融合与临床反馈迭代机制,以提升安全性、可解释性与人机协作水平。尽管无法独立承担儿科诊疗,它们在辅助决策、医学教育与患者沟通方面具有潜力,为构建安全可信的智能儿科医疗系统奠定基础。

原文摘要 · Abstract (English)

With the rapid rise of large language models (LLMs) in medicine, a key question is whether they can function as competent pediatricians in real-world clinical settings. We developed PEDIASBench, a systematic evaluation framework centered on a knowledge-system framework and tailored to realistic clinical environments. PEDIASBench assesses LLMs across three dimensions: application of basic knowledge, dynamic diagnosis and treatment capability, and pediatric medical safety and medical ethics. We evaluated 12 representative models released over the past two years, including GPT-4o, Qwen3-235B-A22B, and DeepSeek-V3, covering 19 pediatric subspecialties and 211 prototypical diseases. State-of-the-art models performed well on foundational knowledge, with Qwen3-235B-A22B achieving over 90% accuracy on licensing-level questions, but performance declined ~15% as task complexity increased, revealing limitations in complex reasoning. Multiple-choice assessments highlighted weaknesses in integrative reasoning and knowledge recall. In dynamic diagnosis and treatment scenarios, DeepSeek-R1 scored highest in case reasoning (mean 0.58), yet most models struggled to adapt to real-time patient changes. On pediatric medical ethics and safety tasks, Qwen2.5-72B performed best (accuracy 92.05%), though humanistic sensitivity remained limited. These findings indicate that pediatric LLMs are constrained by limited dynamic decision-making and underdeveloped humanistic care. Future development should focus on multimodal integration and a clinical feedback-model iteration loop to enhance safety, interpretability, and human-AI collaboration. While current LLMs cannot independently perform pediatric care, they hold promise for decision support, medical education, and patient communication, laying the groundwork for a safe, trustworthy, and collaborative intelligent pediatric healthcare system.

大模型儿科医疗AI评估人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。