arXiv:2606.11208cs.CLcs.AI2026-06

为医学论文中的表面矛盾提供分类与解释框架,区分真实冲突与情境差异。

BioDivergence: A Benchmark and Evaluation Framework for Hidden Contextual Contradictions in Biomedical Abstracts

  • 构建六类冲突分类与十三维分化维度,系统解析研究差异根源。
  • 在1.18万组论文对上验证,模型准确率达55.23%,上下文F1达38.94%。
  • 适合医学AI、可解释性研究者使用,推动更精准的科学结论比对。

医学研究常看似相互矛盾,实则多由人群、地理、检测方法、疾病亚型及临床环境等情境因素导致,双方结论在各自条件下均成立。现有自然语言推理与科学论断验证基准将此类情况简化为蕴含、矛盾或中立,未能捕捉其背后的情境结构。为此,我们提出BioDivergence,包含六类冲突分类、十三轴分化本体,以及每对论断的四类结构化输出:冲突类型、分化轴、主导混杂因子和调和解释。发布BioDivergence-Silver-v1.0,一个跨五大学科领域的11,865组论断对文章无关银标准数据集,并提供去重历史版本用于对比。结果表明两者排名差异显著,微调参考模型在文章无关设定下性能下降约12分;Mistral-7B-Instruct-v0.3在842例主测试集上取得0.5523准确率与0.3894上下文F1。BioDivergence为区分情境性分歧与直接矛盾提供了更忠实的评估方式,并有效分离文章级记忆与真正任务学习。

原文摘要 · Abstract (English)

Biomedical findings often seem to conflict across studies, but many of these differences are context-dependent rather than true contradictions. Variations in cohort, geography, assay protocol, disease subtype, and clinical setting can make both claims locally valid. Existing NLI and scientific claim-verification benchmarks reduce such cases to entailment, contradiction, or neutral, failing to capture the contextual structure behind divergence. To address this, we introduce BioDivergence, an evaluation framework with a six-class conflict taxonomy, a 13-axis divergence ontology, and four structured outputs per claim pair: conflict type, divergence axes, dominant confounder, and reconciliation explanation. We release BioDivergence-Silver-v1.0, an article-disjoint silver benchmark of 11,865 claim pairs across five biomedical domains, alongside a legacy deduplicated variant for comparison. Results show notable ranking differences between the two variants, with the fine-tuned reference model dropping about 12 points under the article-disjoint setting, while Mistral-7B-Instruct-v0.3 achieves 0.5523 accuracy and 0.3894 contextual-F1 on the 842-example primary test set. BioDivergence offers a more faithful way to distinguish contextual divergence from direct contradiction and to separate article-level memorization from genuine task learning.

医学AI文本理解可解释性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。