arXiv:2608.18116cs.CLcs.LG2026-08中稿 · the journal Proces…

优化农业视觉语言模型的提示词,提升跨领域识别准确率

You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models

论文配图:You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models
图 1 · 摘自论文原文
  • 用提示词加权策略自动筛选有效提示,适应不同领域
  • 在域外数据上性能提升显著,51-52个专用提示优于426个通用提示
  • 提出新方法检测模型不确定度,适合高风险农业场景应用

视觉语言模型通过自然语言提示实现零样本分类,但性能对提示设计敏感,尤其在农业食品领域。零样本提示集成(ZPE)通过判别信号加权提示,但其在域偏移下的表现尚不明确。我们在四个数据集和四组提示池上评估了CLIP与SigLIP在农业食品领域的表现,涵盖分布内(ID)食物和分布外农业基准。ZPE在ID条件下改善有限,但在域偏移下显著提升性能与校准能力,特定领域提示池(51-52个)始终优于通用提示池(247-426个)。词汇分析表明,ZPE可无监督地检测域偏移。我们进一步提出基于提示不一致性的检测方法(PID),利用提示分歧表示认知不确定性,在严重域偏移下优于传统置信度指标。

原文摘要 · Abstract (English)

Vision-language models enable zero-shot classification through natural language prompts, but performance is sensitive to prompt formulation, especially in specialized domains. Zero-shot Prompt Ensembling (ZPE) addresses this by weighting prompts by discriminative signal, yet its behavior under domain shift remains unexplored. We evaluate ZPE in the agrifood domain using CLIP and SigLIP across four datasets and four prompt pools, spanning in-distribution (ID) food and out-of-distribution agricultural benchmarks. ZPE provides limited benefit under ID conditions but substantially improves performance and calibration under domain shift, where domain-specific pools of 51-52 prompts consistently outperform generic pools of 247-426. Lexical analysis shows that ZPE acts as an unsupervised domain-alignment detector without label access. We further introduce PID (Prompt-based Inconsistency Detection), which repurposes prompt disagreement as epistemic uncertainty, improving failure detection under severe domain shift where standard confidence measures collapse.

视觉语言模型农业图像提示工程不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。