用心理治疗指南测试大模型推理能力,发现检索到正确资料仍会出错。
CARE-RAG - Clinical Assessment and Reasoning in RAG
- 以临床指南为基准,评估大模型在给定权威文本下的推理质量
- 即使提供正确参考资料,模型仍存在推理错误,准确率未达理想水平
- 提出新评估框架,强调推理一致性与真实性比单纯检索更重要
获取正确的证据并不保证大语言模型(LLMs)能正确推理。在临床场景中,输出必须符合结构化协议,这一“检索与推理之间的鸿沟”尤为严重。我们以书面暴露疗法(WET)指南为测试基准,评估模型对经临床医生筛选问题的回应。结果显示,即便提供了权威文本,错误依然存在。为此,我们提出一个评估框架,衡量推理的准确性、一致性和保真度。结果表明,检索增强生成(RAG)虽能约束输出,但安全部署需像评估检索一样严格评估推理过程。
原文摘要 · Abstract (English)
Access to the right evidence does not guarantee that large language models (LLMs) will reason with it correctly. This gap between retrieval and reasoning is especially concerning in clinical settings, where outputs must align with structured protocols. We study this gap using Written Exposure Therapy (WET) guidelines as a testbed. In evaluating model responses to curated clinician-vetted questions, we find that errors persist even when authoritative passages are provided. To address this, we propose an evaluation framework that measures accuracy, consistency, and fidelity of reasoning. Our results highlight both the potential and the risks: retrieval-augmented generation (RAG) can constrain outputs, but safe deployment requires assessing reasoning as rigorously as retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。