arXiv:2601.03627cs.CLcs.AI2026-01Conference of the …被引 1

用诊断指南评估大模型问诊能力,小模型微调后表现超顶尖模型。

Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines

  • 基于诊断指南直接对比病史与诊疗建议,间接验证疾病诊断能力。
  • 微调的小模型在问诊任务中超越前沿大模型,但病史越长未必诊断越准。
  • 问诊语言风格影响对话特征,适合临床智能助手研发者参考。

我们提出EPAG,一个用于评估大语言模型(LLMs)在诊疗前问诊能力的基准数据集与框架。通过将模型输出与病史采集(HPI)诊断指南进行直接比对,并结合疾病诊断结果进行间接评估,实验发现:经过精心筛选、任务特定数据微调的小型开源模型,在预问诊任务中可超越当前前沿大模型;同时,病史信息量增加并不必然提升诊断准确率。进一步分析显示,问诊所用语言会影响对话特征。我们已将数据集和评估流程开源至https://github.com/seemdog/EPAG,旨在推动真实临床场景中大模型应用的评估与发展。

原文摘要 · Abstract (English)

We introduce EPAG, a benchmark dataset and framework designed for Evaluating the Pre-consultation Ability of LLMs using diagnostic Guidelines. LLMs are evaluated directly through HPI-diagnostic guideline comparison and indirectly through disease diagnosis. In our experiments, we observe that small open-source models fine-tuned with a well-curated, task-specific dataset can outperform frontier LLMs in pre-consultation. Additionally, we find that increased amount of HPI (History of Present Illness) does not necessarily lead to improved diagnostic performance. Further experiments reveal that the language of pre-consultation influences the characteristics of the dialogue. By open-sourcing our dataset and evaluation pipeline on https://github.com/seemdog/EPAG, we aim to contribute to the evaluation and further development of LLM applications in real-world clinical settings.

大模型评估临床问诊诊断指南医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。