测试大模型在医学矛盾信息中的推理能力,发现其能识别正确信息并忽略错误内容。
HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
- 构建920个医学问题对,含相互矛盾的科学文献支持
- 微调模型在正确上下文下准确率超85%,错误上下文干扰显著
- 适合关注医疗AI可信度与抗干扰能力的研究者
语言模型如何利用上下文回答健康问题?其响应受矛盾信息影响如何?我们通过HealthContradict——一个由专家验证的920个独特实例组成的基准数据集,评估语言模型在长篇、相互矛盾的生物医学上下文中的推理能力。每个实例包含一个健康相关问题、一个有科学证据支持的事实答案,以及两份立场相反的文档。我们测试了多种提示设置(正确、错误或矛盾上下文),并测量其对模型输出的影响。相比现有医学问答评估基准,HealthContradict更清晰地区分了模型的上下文推理能力。实验表明,经过微调的生物医学语言模型的能力不仅源于预训练时的参数化知识,更体现在其能够利用正确上下文同时抵制错误上下文干扰的能力。
原文摘要 · Abstract (English)
How do language models use contextual information to answer health questions? How are their responses impacted by conflicting contexts? We assess the ability of language models to reason over long, conflicting biomedical contexts using HealthContradict, an expert-verified dataset comprising 920 unique instances, each consisting of a health-related question, a factual answer supported by scientific evidence, and two documents presenting contradictory stances. We consider several prompt settings, including correct, incorrect or contradictory context, and measure their impact on model outputs. Compared to existing medical question-answering evaluation benchmarks, HealthContradict provides greater distinctions of language models' contextual reasoning capabilities. Our experiments show that the strength of fine-tuned biomedical language models lies not only in their parametric knowledge from pretraining, but also in their ability to exploit correct context while resisting incorrect context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。