arXiv:2606.20873cs.CL2026-06

用结构化推理验证科学论断,提升准确率与可解释性。

SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding

论文配图:SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding
图 1 · 摘自论文原文
  • 将论断分解为原子单元,分模态匹配证据
  • 在SciClaimEval上达79.2%宏F1,63.1%配对准确率
  • 适合需要高可信度验证的科研自动化场景

科学发现日益依赖自动化系统生成假设、检视多模态证据并大规模验证论断。然而,直接让视觉语言模型做二元判断难以满足需求:论断常包含数值结果、比较关系、范围限定和解释背景,而证据则以表格和图表形式存在,具有特定的定位结构。本文提出SciLens,一种基于证据条件的原子蕴含框架,用于多模态科学论断验证。SciLens将每个论断分解为核心实证原子与背景原子,将核心原子分别关联到表格或图表中的具体证据单元(如表格的行、列、单元格、算术关系、表域;图表的面板、坐标轴、图例、视觉编码、类别、趋势、排名、限定词检查),并通过原子级蕴含规则预测最终标签。只有当所有核心实证原子均被当前证据支持时,论断才被视为成立。在SciClaimEval开发集上,SciLens达到79.2%宏F1和63.1%配对准确率,表明结构化的代理式验证能有效提升证据敏感性和可解释性。

原文摘要 · Abstract (English)

Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification is not well served by asking a vision-language model for a direct binary judgment: claims often combine numerical results, comparisons, scope qualifiers, and explanatory context, while evidence is encoded in tables and figures with distinct grounding structures. We present SciLens, an evidence-conditioned atomic entailment framework for multimodal scientific claim verification. SciLens decomposes each claim into central empirical atoms and background atoms, grounds the central atoms to modality-specific evidence witnesses, and predicts the final label with an atom-level entailment rule. For tables, atoms are grounded to rows, columns, cells, arithmetic relations, and table scope; for figures, they are grounded through panels, axes, legends, visual encodings, categories, trends, ranks, and qualifier checks. This yields a unified validation procedure in which a claim is supported only if every central empirical atom is entailed by the current evidence. On the SciClaimEval development set, SciLens achieves 79.2% macro-F1 and 63.1% pair accuracy, showing that structured agentic validation improves both evidence sensitivity and interpretability.

科学验证多模态逻辑推理证据锚定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。