arXiv:2604.12403cs.CV2026-04中稿 · CVPR被引 1

用图文双锚点筛选有效图像,提升测试时提示调优效果

Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning

论文配图:Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning
图 1 · 摘自论文原文
  • 引入文本与图像双锚点,基于语义对齐筛选有用视图
  • 在15个数据集上达新最优,显著提升模型鲁棒性
  • 适合需要稳定微调的视觉语言模型应用

测试时提示调优(TPT)通过增强视图适配视觉-语言模型,但难以判断哪些视图有益。传统基于熵的过滤依赖模型内部置信度,分布偏移下常误判无关区域为高置信,忽略语义内容。为此,我们提出双模态锚点引导框架,将视图选择锚定于语义证据:引入属性丰富的文本锚点提供细粒度类别语义,设计自适应图像锚点捕捉测试时统计演化。基于对齐与置信度过滤视图,确保仅信息量高的视图指导适应。此外,将锚点视为辅助预测头,以置信度加权融合其输出与原模型,生成稳定监督信号用于提示更新。在15个基准数据集上的大量实验表明性能达到新最佳,验证了锚点引导监督作为鲁棒提示更新基础的有效性。

原文摘要 · Abstract (English)

Test-Time Prompt Tuning (TPT) adapts vision-language models using augmented views, but its effectiveness is hindered by the challenge of determining which views are beneficial. Standard entropy-based filtering relies on the internal confidence scores of the model, which are often miscalibrated under distribution shift, assigning high confidence to irrelevant crops or background regions while ignoring semantic content. To address this, we propose a dual-modality anchor-guided framework that grounds view selection in semantic evidence. We introduce a text anchor from attribute-rich descriptions, to provide fine-grained class semantics, and an adaptive image anchor that captures evolving test-time statistics. Using these anchors, we filter views based on alignment and confidence, ensuring that only informative views guide adaptation. Moreover, we treat the anchors as auxiliary predictive heads and combine their predictions with the original output in a confidence-weighted ensemble, yielding a stable supervision signal for prompt updates. Extensive experiments on 15 benchmark datasets demonstrate new state-of-the-art performance, highlighting the contribution of anchor-guided supervision as a foundation for robust prompt updates.

提示调优视觉语言多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。