arXiv:2504.02144cs.LGcs.AI2025-04被引 4

提出可解释软提示新框架,揭示性能与可解释性间的根本权衡。

Towards Interpretable Soft Prompts

  • 基于忠实性与可审视性构建可解释性评估理论框架
  • 实验表明现有方法无法自然满足可解释性要求
  • 新目标函数使软提示在保持性能前提下提升可读性

软提示作为低成本提升大模型任务表现的方法已广受欢迎。尽管起源于自动化提示设计,但软提示及其他可训练提示仍是黑箱方法,缺乏与提示语义的直接可解释关联。本文提出一种新的理论框架,基于两个理想属性——忠实性与可审视性,评估可训练提示的可解释性。研究发现,现有方法并不天然满足该可解释性标准。由此启发出一种新方向:显式优化可解释性的可训练提示方法。为此,我们为两种前沿提示调优器(PEZ 和 RLPrompt)设计并测试了新型面向可解释性的目标函数。在 GPT-2 上的实验揭示了可解释性与任务性能之间存在根本权衡,同时暴露了仅优化可解释性代理指标时出现的异常行为。

原文摘要 · Abstract (English)

Soft prompts have been popularized as a cheap and easy way to improve task-specific LLM performance beyond few-shot prompts. Despite their origin as an automated prompting method, however, soft prompts and other trainable prompts remain a black-box method with no immediately interpretable connections to prompting. We create a novel theoretical framework for evaluating the interpretability of trainable prompts based on two desiderata: faithfulness and scrutability. We find that existing methods do not naturally satisfy our proposed interpretability criterion. Instead, our framework inspires a new direction of trainable prompting methods that explicitly optimizes for interpretability. To this end, we formulate and test new interpretability-oriented objective functions for two state-of-the-art prompt tuners: Hard Prompts Made Easy (PEZ) and RLPrompt. Our experiments with GPT-2 demonstrate a fundamental trade-off between interpretability and the task-performance of the trainable prompt, explicating the hardness of the soft prompt interpretability problem and revealing odd behavior that arises when one optimizes for an interpretability proxy.

可解释性软提示提示调优大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。