让提示学习既可解释又高效,通过交替选择语义词与优化连续提示。
Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning

- 交替进行语义词选择和连续提示优化,兼顾可解释性与适应性。
- 在五个方法上提升准确率,同时显著增强提示的可读性。
- 无需外部大模型,适合希望提升可解释性的研究者使用。
视觉语言模型如CLIP在连续提示学习中常出现过拟合和可解释性差的问题。虽然离散提示优化提升可解释性,但通常依赖大型外部模型,导致计算成本高且难以扩展。本文提出可解释提示学习(IPL),一种混合框架,交替进行离散语义词选择与连续提示优化。具体地,将语义词选择建模为近似子模优化问题,鼓励人类可理解且语义多样的词汇。同时采用交替优化策略,融合离散选择与连续调优,在保持下游任务适应性的同时提升可解释性。该框架即插即用,可无缝集成到现有提示学习方法中。在多个基准上的实验表明,IPL在五个代表性提示学习方法中均一致提升准确率与可解释性,为现有框架提供了有效且可扩展的延伸。
原文摘要 · Abstract (English)
Vision-language models such as CLIP achieve strong visual-textual alignment, but often suffer from overfitting and limited interpretability when adapted through continuous prompt learning. While discrete prompt optimization improves interpretability, it usually depends on large external models, leading to high computational costs and limited scalability. In this paper, we propose Interpretable Prompt Learning (IPL), a hybrid framework that alternates between discrete semantic token selection and continuous prompt optimization. Specifically, IPL formulates semantic token selection as an approximate submodular optimization problem, encouraging tokens that are both human-understandable and semantically diverse. It further adopts an alternating optimization strategy to integrate discrete token selection with continuous prompt tuning, improving interpretability while preserving adaptability to downstream tasks. Our framework is plug-and-play, allowing seamless integration with existing prompt learning methods. Extensive experiments on multiple benchmarks show that IPL consistently improves both interpretability and accuracy across five representative prompt learning methods, providing an effective and scalable extension to existing frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。