arXiv:2605.18419cs.CVcs.AI2026-05

提出无需训练的几何感知核心集,提升病理图像少样本学习的稳定性和可靠性。

Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology

论文配图:Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology
图 1 · 摘自论文原文
  • 在预训练多模态嵌入空间中,联合优化分布保真、语义对齐与不确定性抑制。
  • 在多个模型和数据集上达到最强基线精度,同时显著改善校准与抗提示扰动能力。
  • 适合追求高可靠性的医学视觉语言模型应用,尤其适用于标注稀缺场景。

视觉语言模型(VLM)可结合视觉感知与开放性临床推理,适用于计算病理学。然而,在稀少且专家标注的病理数据上微调数十亿参数代价高昂;而无需参数更新的上下文学习(ICL)对示例选择和查询表述高度敏感,导致诊断不可靠。现有选择策略依赖查询相关的最近邻检索,忽略全局数据结构,需昂贵的参数更新,或忽视VLM的联合视文本嵌入几何特性。本文提出GAUC,一种直接在预训练多模态嵌入空间中运行的免训练核心集选择方法。GAUC联合优化三项目标:(1) 最大均值差异项,确保核心集与全数据集的分布保真;(2) 有效互信息差正则项,通过利用VLM的联合视文本对齐,限制提示改写下的性能退化;(3) 预测不确定性(熵)惩罚项,抑制模糊、易幻觉的输出。在CRC-100K和MHIST数据集上,针对多种开源VLM架构,GAUC在准确率上匹配最强的ICL选择与数据蒸馏基线,同时大幅改善校准度、提示鲁棒性及幻觉率,且全程无需梯度更新。

原文摘要 · Abstract (English)

Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology. However, fine-tuning billions of parameters on scarce, expert-annotated pathology data is prohibitive, while in-context learning (ICL), which conditions the VLM on demonstrative image-text pairs without parameter updates, suffers from high sensitivity to which examples are selected and how the query is phrased, producing unreliable diagnostics. Existing selection strategies rely on query-dependent nearest-neighbour retrieval that ignores global data structure, require costly parameter updates, or disregard the joint vision-text embedding geometry of VLMs. We propose GAUC, a training-free coreset selection method operating directly in the pre-trained multimodal embedding space. GAUC jointly optimises three objectives: (1) a Maximum Mean Discrepancy term enforcing distributional fidelity between coreset and full dataset, (2) an Effective Mutual Information Difference regulariser bounding performance degradation under prompt paraphrases by exploiting the VLM's joint vision-text alignment, and (3) a predictive-uncertainty (entropy) penalty suppressing ambivalent, hallucination-prone outputs. On CRC-100K and MHIST across multiple open-source VLM architectures, GAUC \emph{matches} the accuracy of the strongest ICL selection and dataset-distillation baselines while substantially improving calibration, prompt robustness, and hallucination rates, all without a single gradient update.

医学图像少样本学习视觉语言模型不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。