让医学问答更可信:用结构化证据和不确定性融合提升答案可靠性
BELIEF: Structured Evidence Modeling and Uncertainty-Aware Fusion for Biomedical Question Answering

- 将文献转化为含属性的证据对象,显式记录来源质量与支持强度
- 符号与神经双路径融合,使不确定性和证据冲突可见,准确率最高提升25%
- 适合需要高可信度医学决策的场景,尤其适合通用大模型落地
医学问答常需从检索文献中做出判断,但文献的相关性、质量及对答案的支持度参差不齐。现有检索增强型大模型多将文献作为扁平文本输入,导致证据可靠性与剩余不确定性隐含不清。我们提出BELIEF框架,通过结构化证据建模与不确定性感知融合,实现闭集医学问答。该框架将检索文档转化为包含临床属性、源质量、问题相关性、支持强度及对应假设的证据对象,为符号与神经双路径提供统一基础。符号路径基于有限答案空间,采用Dempster–Shafer理论构建可靠性加权基本概率分配,实现不确定性感知的符号融合以估计信念与残余不确定性;神经路径则利用相同结构化证据进行大模型语义推理,再通过可靠性感知仲裁模块根据信念强度、不确定性、证据可靠性与语义一致性协调两路输出。在PubMedQA、MedQA和MedMCQA三个数据集上,使用五种通用大模型骨干网络的实验表明,贝尔在30组设置中取得最佳结果25次。与领域专用模型对比,贝尔在MedQA和MedMCQA上表现竞争力,但在PubMedQA上仍逊于专门预训练模型。消融实验、互补性分析、不确定性分层分析与成本分析进一步验证,贝尔通过显式表达证据结构、路径分歧与决策不确定性,显著提升了检索证据的利用率。
原文摘要 · Abstract (English)
Biomedical question answering often requires decisions from retrieved literature whose relevance, quality, and support for candidate answers are uneven. Most retrieval-augmented large language model (LLM) methods feed this literature to the model as flat text, leaving evidence reliability and remaining uncertainty largely implicit. We propose BELIEF, a structured evidence modeling and uncertainty-aware fusion framework for closed-set biomedical question answering. Rather than treating retrieved documents as undifferentiated context, BELIEF converts them into evidence objects that record clinical attributes, source quality, question relevance, support strength, and the associated candidate hypothesis. These evidence objects provide a shared basis for two complementary reasoning paths. The symbolic path constructs reliability-weighted basic probability assignments based on Dempster--Shafer (D-S) theory over a finite answer space and performs uncertainty-aware symbolic evidence fusion to estimate belief and residual uncertainty. The neural path uses the same structured evidence for LLM-based semantic inference, while a reliability-aware arbitration module reconciles the symbolic and neural outputs according to belief strength, uncertainty, evidence reliability, and semantic consistency. Experiments on PubMedQA, MedQA, and MedMCQA with five general-purpose LLM backbones show that BELIEF obtains the best result in 25 of 30 backbone--dataset--metric settings. Comparisons with biomedical-domain models indicate that BELIEF is competitive on MedQA and MedMCQA, while specialized biomedical pretraining remains advantageous on PubMedQA. Ablation, complementarity, uncertainty-stratified, and cost analyses further show that BELIEF improves retrieved-evidence utilization by making evidence structure, path disagreement, and decision uncertainty explicit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。