arXiv:2607.23794cs.CVcs.AI2026-07被引 1

构建抗干扰的跨尺度病理图像理解基准,提升模型多倍率分析能力。

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

论文配图:PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis
图 1 · 摘自论文原文
  • 设计对抗性文本筛查与结构化干扰采样,防止模型依赖表面线索。
  • 构建含10,373题的跨尺度病理VQA数据集,覆盖1,368个诊断路径。
  • 通过难度驱动蒸馏与尺度感知奖励,显著提升多尺度推理能力。

病理诊断本质上是多尺度的,需融合低倍率下的组织整体结构与高倍率下的细胞形态。然而现有病理基准和视觉语言模型(VLMs)仍主要基于单尺度设置,限制了其学习临床有意义的多倍率推理能力。此外,粗略构建的视觉问答(VQA)任务可能受文本或浅层视觉线索影响,导致视觉理解评估不可靠。为此,我们提出一个抗捷径的跨尺度病理推理基准与训练框架。设计对抗性文本仅筛查策略用于语义推理问题,以及结构控制的干扰采样策略用于视觉定位问题,促使模型依赖跨尺度视觉证据。基于此流程,构建了PathScale-VQA——一个高质量的跨尺度病理VQA基准,包含10,373道多选题,基于1,368个诊断路径,覆盖多个放大倍率。在语义推理集基础上,通过难度驱动的推理蒸馏微调及带尺度感知推理结构奖励的强化学习优化,得到PathScale-R1模型。大量实验表明,PathScale-R1在跨尺度推理任务上达到领先性能,并有效迁移至传统单尺度病理VQA任务。

原文摘要 · Abstract (English)

Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) are still largely developed under single-scale settings, limiting their ability to learn clinically meaningful multi-magnification reasoning. Moreover, naively constructed visual question answering (VQA) tasks may be susceptible to text-only or superficial visual shortcuts, leading to unreliable assessments of visual understanding. To address these limitations, we introduce a benchmark and training framework for shortcut-resistant cross-scale pathology reasoning. We design an Adversarial Text-only Screening strategy for semantic reasoning questions and a Structure-controlled Distractor Sampling strategy for visual grounding questions, encouraging models to rely on cross-scale visual evidence. Based on this pipeline, we construct PathScale-VQA, a high-quality cross-scale pathology VQA benchmark with 10,373 multiple-choice questions grounded in 1,368 diagnostic paths across multiple magnification levels. Building on the semantic reasoning set, PathScale-R1 is optimized through Difficulty-driven Reasoning Distillation supervised fine-tuning followed by reinforcement learning with a Scale-aware Reasoning Structure reward, which encourages the use of evidence across magnifications. Extensive experiments demonstrate state-of-the-art performance of PathScale-R1 on cross-scale reasoning tasks and effective transfer to conventional single-scale pathology VQA. Our code is available at https://github.com/iMVR-PL/PathScale-R1.

病理分析多尺度视觉问答模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。