arXiv:2608.17607cs.CV2026-08

测试病理图像推理中证据是否真实支撑答案,发现高准确率模型仍可能不靠谱。

PathoArgus: Advancing Evidence-Grounded Long-Context Visual Reasoning across Gigapixel Whole-Slide and Multi-Slide Case Contexts

论文配图:PathoArgus: Advancing Evidence-Grounded Long-Context Visual Reasoning across Gigapixel Whole-Slide and Multi-Slide Case Contexts
图 1 · 摘自论文原文
  • 设计新评测基准,检验模型是否真从大图中找到证据
  • 仅1.86%的预测在证据变换后保持一致,暴露现有模型依赖文本线索
  • 适合关注医学AI可解释性与真实推理能力的研究者

全切片病理分析需整合跨完整病例的亿级像素视觉证据,但现有问答基准仅衡量最终答案准确率,易受语言先验干扰且无法验证预测是否基于实际组织切片。我们提出PathoArgus-Bench,包含4,913名患者的22,078道四选题,覆盖TCGA中15个项目的六种病理能力,分三个证据需求层级,并设定固定阅读预算,仅保留极小部分原始上下文。为隔离证据依赖性,引入ESG(证据状态四联体):问题文本固定,目标WSI集被移动、替换或移除,要求预测一致。评估20个通用、医学及病理专用系统发现:即使GPT-5.6达到57.09%总体准确率与57.04%ESG准确率,其在483个四联体中仅正确完成19个(QExact 3.93%),表明行级准确率不能代表可靠证据支撑。我们还提出PathoArgus,一种基于问题相关性与空间覆盖率分配上下文的固定预算阅读器,达50.39%准确率,但QExact仅为1.86%,证明提升上下文获取并不等于实现稳定证据推理。该基准与诊断工具表明,获取有用上下文是必要前提,却远不充分,亟需从以答案为中心转向以证据为基础的评估范式。

原文摘要 · Abstract (English)

Whole-slide pathology reasoning requires models to integrate gigapixel-scale visual evidence across complete case-linked slides, yet current question-answering benchmarks primarily measure final answer accuracy--a metric vulnerable to linguistic priors and benchmark regularities, and insufficient to establish that predictions are grounded in the supplied tissue. We introduce PathoArgus-Bench, a benchmark and evaluation protocol that explicitly tests the full evidence chain: availability, accessibility, use, and responsiveness. PathoArgus-Bench comprises 22,078 four-choice questions from 4,913 patients across 15 TCGA projects, covering six pathology capabilities across three levels of evidence demand, and operates under a fixed reader budget that retains only a small fraction of the gigapixel context. To further isolate evidence-grounded reasoning, we contribute ESG (Evidence State Quartets), a controlled set of 483 quartets where the question text is fixed while the target WSI set is moved, replaced, or removed, requiring consistent predictions across all states. Evaluating 20 general-purpose, medical, and pathology-specific systems reveals a stark gap: while GPT-5.6 achieves 57.09% overall accuracy and 57.04% on ESG, it correctly completes only 19 of 483 quartets (3.93% QExact), exposing that row-level accuracy does not translate into reliable evidence grounding. We also introduce PathoArgus, a fixed-budget reader that allocates context via question relevance and spatial coverage, attaining 50.39% overall accuracy yet only 1.86% QExact--demonstrating that improved context access alone does not ensure consistent evidence-based prediction. Our benchmark and diagnostics establish that acquiring useful whole-slide context is necessary but far from sufficient, and call for a shift from answer-centric to evidence-grounded evaluation in computational pathology.

病理分析证据推理大图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。