arXiv:2507.18143cs.CLcs.AI2025-07被引 1

评测大模型在艾滋病诊疗中的表现,发现顶级模型仍存认知偏差。

HIVMedQA: Benchmarking large language models for HIV medical decision support

  • 构建艾滋病临床问答基准HIVMedQA,融合医生经验设计问题
  • Gemini 2.5 Pro表现最佳,但复杂问题下性能明显下降
  • 专业微调模型未必更优,大模型不等于高可靠,适合临床评估者参考

大型语言模型(LLMs)正成为临床决策支持的新兴工具。艾滋病管理因其治疗方案多样、合并症复杂及依从性挑战,是理想的应用场景。然而,将其融入临床实践面临准确性、潜在伤害和医生接受度等担忧。尽管前景广阔,当前针对艾滋病护理的AI应用研究仍较少,且缺乏对LLM的系统性评估。本研究评估了现有LLM在艾滋病管理中的能力,揭示其优势与局限。我们提出HIVMedQA,一个专为评估艾滋病临床开放问答设计的基准数据集,问题由感染病医生共同筛选并优化。测试涵盖7个通用型和3个医学专用型模型,通过提示工程提升表现。评估框架结合词汇相似性与“大模型作为评判者”方法,并扩展以更贴近临床实际。评估维度包括问题理解、推理、知识回忆、偏见、潜在危害和事实准确性。结果显示,Gemini 2.5 Pro在多数维度中表现领先;其中两个排名前三位的模型为专有模型。随着问题复杂度上升,性能显著下降。医学微调模型并非始终优于通用模型,模型规模也非性能可靠预测指标。推理与理解难度高于事实记忆,观察到近期效应与现状偏倚等认知偏差。研究强调需针对性开发与评估,以确保大模型在临床应用中的安全与有效。

原文摘要 · Abstract (English)

Large language models (LLMs) are emerging as valuable tools to support clinicians in routine decision-making. HIV management is a compelling use case due to its complexity, including diverse treatment options, comorbidities, and adherence challenges. However, integrating LLMs into clinical practice raises concerns about accuracy, potential harm, and clinician acceptance. Despite their promise, AI applications in HIV care remain underexplored, and LLM benchmarking studies are scarce. This study evaluates the current capabilities of LLMs in HIV management, highlighting their strengths and limitations. We introduce HIVMedQA, a benchmark designed to assess open-ended medical question answering in HIV care. The dataset consists of curated, clinically relevant questions developed with input from an infectious disease physician. We evaluated seven general-purpose and three medically specialized LLMs, applying prompt engineering to enhance performance. Our evaluation framework incorporates both lexical similarity and an LLM-as-a-judge approach, extended to better reflect clinical relevance. We assessed performance across key dimensions: question comprehension, reasoning, knowledge recall, bias, potential harm, and factual accuracy. Results show that Gemini 2.5 Pro consistently outperformed other models across most dimensions. Notably, two of the top three models were proprietary. Performance declined as question complexity increased. Medically fine-tuned models did not always outperform general-purpose ones, and larger model size was not a reliable predictor of performance. Reasoning and comprehension were more challenging than factual recall, and cognitive biases such as recency and status quo were observed. These findings underscore the need for targeted development and evaluation to ensure safe, effective LLM integration in clinical care.

大模型评估艾滋病诊疗医疗AI认知偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。