arXiv:2511.19652cs.CV2025-11被引 6

用通用模型自主导航病理大图,实现多尺度精准问答。

Navigating Gigapixel Pathology Images with Large Multimodal Models

  • 无需训练,通过迭代选择不同倍率图像块自主探索病理大片。
  • 在5个临床任务中4项达当前最优,多尺度推理能力显著提升。
  • 适合医学影像研究者、病理AI开发者快速构建可泛化系统。

大型多模态模型的进展使得交互式聊天系统能够对病理全幻灯片图像(WSIs)进行对话和推理。然而,现有基于幻灯片级别的聊天系统通常高度专用,通常将WSIs压缩为固定幻灯片级嵌入或依赖多组件流水线,这会丢失多尺度细节并限制在目标任务之外的泛化能力。我们提出GIANT(Gigapixel Image Agent for Navigating Tissue),一种简单且无需训练的方法,使通用多模态模型能自主导航WSIs,迭代选择多倍率图像区域并随时间聚合证据。为评估在WSI问答中的泛化能力并促进可复现性,我们引入MultiPathQA,一个涵盖五个临床挑战、覆盖868张独特WSIs的基准套件,包含934个问题,其中128个由病理科医生编写的多选题,旨在模拟真实的诊断搜索与多尺度推理。使用GPT-5时,GIANT在五个基准中的四项上表现超越专用于病理问答的模型,达到当前最优水平。

原文摘要 · Abstract (English)

Recent advances in large multimodal models have allowed for the development of interactive chat models that can converse and reason about pathology whole-slide images (WSIs). However, existing slide-level chat systems are often highly specialized, typically compressing WSIs into fixed slide-level embeddings or relying on multi-component pipelines, which can lose multi-scale detail and limit generalizability beyond the target task. We present GIANT (Gigapixel Image Agent for Navigating Tissue), a simple, training-free approach that lets general-purpose multimodal models navigate WSIs on their own, iteratively selecting multi-magnification crops and aggregating evidence over time. To evaluate generalizability in WSI question answering and to promote reproducibility, we introduce MultiPathQA, a benchmark suite spanning five clinical challenges and 934 questions over 868 unique WSIs. This includes a new set of 128 pathologist-authored multiple-choice questions designed to mirror real diagnostic search and multi-scale reasoning. Using GPT-5, GIANT outperforms models specialized for pathology question answering, achieving state-of-the-art performance on four out of five benchmarks.

病理图像多模态大模型问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。