arXiv:2608.30393cs.CL2026-08

让生物医学AI提取可验证的量化证据,避免虚假可信的结论。

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI

  • 从文献中提取剂量、效应值等量化信息,构建可审计的证据单元
  • 确保每条结论有来源、单位、不确定性及生物学合理性支持
  • 适合医疗决策支持和知识图谱构建者使用

生物医学人工智能系统越来越多地从文献、临床试验和监管文件中提取、组织和重用科学主张。但自动提取并不等于可靠证据:只有当一个主张能追溯到其来源、关联支撑它的量化细节,并置于其生物医学背景和不确定性中时,才真正有用。这一点尤为重要,因为大语言模型(LLMs)和日益自主的系统正推动证据综合、知识图谱(KG)构建与决策支持。当前许多文本挖掘和LLM流程仍以关系为中心,仅捕捉如药物-治疗-疾病这类实体关系,却丢失了剂量、效应大小、人群、对照组、不确定性以及主张成立的条件。此类关系看似可行,实则难以验证、比较或复用。本文倡导转向定量证据挖掘——结构化提取数值、单位、测量实体与属性、上下文、不确定性、出处及可信度,构成可被验证的证据单元,填充可信证据知识图谱。我们提出一种可信赖性感知的AI框架,将提取的主张视为可审计的证据对象,明确说明测量内容、变化量、情境、不确定性及来源。核心风险不仅是错误提取,更是那些看似证据却缺乏可信结构的主张。

原文摘要 · Abstract (English)

Biomedical artificial intelligence (AI) systems increasingly extract, organize, and reuse scientific claims from literature, clinical trials, and regulatory documents. But automatic extraction alone does not make a claim reliable evidence: a claim becomes useful only when it can be traced to its source, linked to the quantitative details that support it, and read within its biomedical context and uncertainty. This matters as large language models (LLMs) and increasingly autonomous systems drive evidence synthesis, knowledge graph (KG) construction, and decision support. Many text-mining and LLM pipelines remain relation-centric: they capture entities and relations such as Drug--TREATS--Disease, but drop the dose, effect size, population, comparator, uncertainty, and conditions under which a claim holds. Such relations can look actionable yet remain hard to verify, compare, or reuse. In this perspective, we argue for a shift toward quantitative evidence mining---extracting values, units, measured entities and properties, context, uncertainty, provenance, and plausibility as structured evidence units that populate evidence-aware KGs and can be checked for source grounding, unit consistency, completeness, and biological plausibility. We outline a framework for plausibility-aware AI that treats extracted claims not as final answers but as auditable evidence objects, making clear what was measured, how much it changed, in which setting, with what uncertainty, and from which source. The central risk is not only incorrect extraction, but claims that look like evidence while lacking the structure needed to trust them.

生物医学AI证据挖掘可解释性知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。