arXiv:2410.14399cs.CL2024-10被引 23

用生物医学本体构建逻辑推理题,测试大模型在医学推理中的表现。

SylloBio-NLI: Evaluating Large Language Models on Biomedical Syllogistic Reasoning

  • 基于生物本体生成28种医学逻辑推理题,评估模型推理能力。
  • 零样本下模型准确率从23%到70%,少数样本提示可提升43%。
  • 模型对词语微小变化敏感,现有模型尚难用于安全医学推断。

命题推理对自然语言推断(NLI)至关重要,在生物医学等专业领域尤其重要,可用于自动证据解析与科学发现。本文提出SylloBio-NLI框架,利用外部本体系统化生成多样化的生物医学命题推理实例,用于评估大语言模型(LLM)在识别有效结论和提取支持证据方面的能力。我们使用该框架,在人类基因组通路基础上的28种命题推理模式上对多个模型进行测试。实验表明,零样本条件下,不同模型在泛化肯定式(modus ponens)上的平均准确率为70%,而在选言推理(disjunctive syllogism)上仅为23%。同时发现,少量示例提示可显著提升性能,如Gemma提升14%,LLama-3提升43%。然而深入分析显示,两种方法均对表面词汇变化高度敏感,表现出可靠性与模型架构、预训练方式之间的依赖性。总体而言,尽管上下文示例有潜力激发模型的命题推理能力,但当前模型仍远未达到安全应用于生物医学推断所需的鲁棒性和一致性。

原文摘要 · Abstract (English)

Syllogistic reasoning is crucial for Natural Language Inference (NLI). This capability is particularly significant in specialized domains such as biomedicine, where it can support automatic evidence interpretation and scientific discovery. This paper presents SylloBio-NLI, a novel framework that leverages external ontologies to systematically instantiate diverse syllogistic arguments for biomedical NLI. We employ SylloBio-NLI to evaluate Large Language Models (LLMs) on identifying valid conclusions and extracting supporting evidence across 28 syllogistic schemes instantiated with human genome pathways. Extensive experiments reveal that biomedical syllogistic reasoning is particularly challenging for zero-shot LLMs, which achieve an average accuracy between 70% on generalized modus ponens and 23% on disjunctive syllogism. At the same time, we found that few-shot prompting can boost the performance of different LLMs, including Gemma (+14%) and LLama-3 (+43%). However, a deeper analysis shows that both techniques exhibit high sensitivity to superficial lexical variations, highlighting a dependency between reliability, models' architecture, and pre-training regime. Overall, our results indicate that, while in-context examples have the potential to elicit syllogistic reasoning in LLMs, existing models are still far from achieving the robustness and consistency required for safe biomedical NLI applications.

逻辑推理生物医学大模型评测知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。