arXiv:2601.17405cs.CV2026-01被引 1

用分层对齐提升大模型在少量样本下发现病理异常的能力

HAAF: Hierarchical Adaptation and Alignment of Foundation Models for Few-Shot Pathology Anomaly Detection

  • 通过跨层级对齐机制,让文本提示与局部图像上下文动态融合
  • 在四个数据集上显著优于现有方法,少样本下表现更稳定
  • 适合医学图像分析、小样本病理诊断等专业领域研究者

精准病理诊断依赖于在特定兴趣区域(ROIs)中识别细微的形态学异常,这些局部纹理线索而非整体切片上下文才是专家诊断的核心依据。尽管视觉-语言(V-L)模型可通过语义先验实现数据高效,但其通用表征难以捕捉此类细微缺陷,存在粒度不匹配问题。现有方法多将模态独立处理,未能将语义提示锚定在ROI级视觉上下文中。为此,本文提出分层自适应与对齐框架(HAAF),核心是跨层级缩放对齐(CLSA)机制:先由视觉特征向文本提示注入上下文,生成内容自适应描述符,再以该描述符空间引导视觉编码器聚焦异常。此外,双分支推理策略融合语义评分与几何原型,增强少样本场景下的稳定性。在四个基准上的实验表明,HAAF显著优于当前最优方法,并可有效适配领域专用骨干网络(如CONCH),在低资源条件下表现优异。

原文摘要 · Abstract (English)

Precision pathology relies on detecting fine-grained morphological abnormalities within specific Regions of Interest (ROIs), as these local, texture-rich cues - rather than global slide contexts - drive expert diagnostic reasoning. While Vision-Language (V-L) models promise data efficiency by leveraging semantic priors, adapting them faces a critical Granularity Mismatch, where generic representations fail to resolve such subtle defects. Current adaptation methods often treat modalities as independent streams, failing to ground semantic prompts in ROI-specific visual contexts. To bridge this gap, we propose the Hierarchical Adaptation and Alignment Framework (HAAF). At its core is a novel Cross-Level Scaled Alignment (CLSA) mechanism that enforces a sequential calibration order: visual features first inject context into text prompts to generate content-adaptive descriptors, which then spatially guide the visual encoder to spotlight anomalies. Additionally, a dual-branch inference strategy integrates semantic scores with geometric prototypes to ensure stability in few-shot settings. Experiments on four benchmarks show HAAF significantly outperforms state-of-the-art methods and effectively scales with domain-specific backbones (e.g., CONCH) in low-resource scenarios.

病理分析少样本学习视觉语言模型图像对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。