arXiv:2409.19487cs.CLcs.LG2024-09被引 51

评估大模型在医疗对话中提问能力的新框架,提升问诊信息获取效率。

HealthQ: Unveiling Questioning Capabilities of LLM Chains in Healthcare Conversations

  • 构建多链式LLM框架,结合RAG、CoT与反思机制生成高质量问题。
  • 通过人工与自动指标验证,优质提问可显著提升患者信息采集完整度。
  • 适用于医疗AI研发者,尤其关注问诊逻辑与交互设计的团队。

数字医疗中有效患者照护依赖于不仅能回答问题,还能通过精心设计的提问主动获取关键信息的大语言模型(LLMs)。本文提出HealthQ,一个评估医疗对话中LLM链提问能力的新型框架。通过实现包含检索增强生成(RAG)、思维链(CoT)和反思链在内的先进LLM链,HealthQ评估其在获取全面且相关患者信息方面的表现。为此,我们引入一个LLM评判器,从具体性、相关性和实用性等维度评估生成问题,并与传统的自然语言处理指标如ROUGE和基于命名实体识别(NER)的集合对比相衔接。我们在两个从公开医学数据集ChatDoctor和MTS-Dialog构建的自定义数据集上验证了HealthQ的有效性,证明其在多个LLM评判模型(包括GPT-3.5、GPT-4和Claude)下均具鲁棒性。本研究贡献三方面:首次系统化提出评估医疗对话中提问能力的框架;建立模型无关的评估方法;提供实证证据,证明高质量提问能提升患者信息获取效果。

原文摘要 · Abstract (English)

Effective patient care in digital healthcare requires large language models (LLMs) that not only answer questions but also actively gather critical information through well-crafted inquiries. This paper introduces HealthQ, a novel framework for evaluating the questioning capabilities of LLM healthcare chains. By implementing advanced LLM chains, including Retrieval-Augmented Generation (RAG), Chain of Thought (CoT), and reflective chains, HealthQ assesses how effectively these chains elicit comprehensive and relevant patient information. To achieve this, we integrate an LLM judge to evaluate generated questions across metrics such as specificity, relevance, and usefulness, while aligning these evaluations with traditional Natural Language Processing (NLP) metrics like ROUGE and Named Entity Recognition (NER)-based set comparisons. We validate HealthQ using two custom datasets constructed from public medical datasets, ChatDoctor and MTS-Dialog, and demonstrate its robustness across multiple LLM judge models, including GPT-3.5, GPT-4, and Claude. Our contributions are threefold: we present the first systematic framework for assessing questioning capabilities in healthcare conversations, establish a model-agnostic evaluation methodology, and provide empirical evidence linking high-quality questions to improved patient information elicitation.

医疗AI对话系统大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。