arXiv:2601.10108cs.CLcs.AI2026-01被引 5

评测大模型理解长篇科学文献能力,要求给出可追溯的跨模态证据链。

SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature

  • 构建跨文本与图表的原生证据链评估框架
  • 多任务测试显示模型答案正确但证据支撑弱
  • 强调可验证证据质量,避免无据之答

评估多模态大模型对长篇科学论文的真实理解能力仍具挑战:仅依赖答案匹配的指标和合成的‘大海捞针’测试常奖励表面正确但缺乏因果证据链的答案。我们提出‘海中鱼’(FITO)范式,要求模型在原始科学文档内构建显式的跨模态证据链。为实现该范式,我们构建了SIN-Data——一个保持原文本与图表自然交错结构的科学文献语料库,并在此基础上构建SIN-Bench,包含四个递进任务:证据发现(SIN-Find)、假设验证(SIN-Verify)、基于证据的问答(SIN-QA)与证据锚定的综合(SIN-Summary)。我们引入“无证据,无得分”评分机制,仅对可验证锚点的预测打分,并通过匹配度、相关性与逻辑性诊断证据质量。在八种MLLM上的实验表明,可追溯性是主要瓶颈:Gemini-3-pro平均得分最高(0.573),而GPT-5虽在SIN-QA上达到0.767的准确率,但在整体证据对齐得分上表现不佳,暴露出正确性与可追踪支持之间的差距。

原文摘要 · Abstract (English)

Evaluating whether multimodal large language models truly understand long-form scientific papers remains challenging: answer-only metrics and synthetic "Needle-In-A-Haystack" tests often reward answer matching without requiring a causal, evidence-linked reasoning trace in the document. We propose the "Fish-in-the-Ocean" (FITO) paradigm, which requires models to construct explicit cross-modal evidence chains within native scientific documents. To operationalize FITO, we build SIN-Data, a scientific interleaved corpus that preserves the native interleaving of text and figures. On top of it, we construct SIN-Bench with four progressive tasks covering evidence discovery (SIN-Find), hypothesis verification (SIN-Verify), grounded QA (SIN-QA), and evidence-anchored synthesis (SIN-Summary). We further introduce "No Evidence, No Score", scoring predictions when grounded to verifiable anchors and diagnosing evidence quality via matching, relevance, and logic. Experiments on eight MLLMs show that grounding is the primary bottleneck: Gemini-3-pro achieves the best average overall score (0.573), while GPT-5 attains the highest SIN-QA answer accuracy (0.767) but underperforms on evidence-aligned overall scores, exposing a gap between correctness and traceable support.

多模态科学文献证据链评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。