arXiv:2505.14943cs.LGcs.AI2025-05

用软提示量化模型行为差距,发现潜在风险能力。

Soft Prompts for Evaluation: Measuring Conditional Distance of Capabilities

  • 用优化输入嵌入(软提示)衡量模型与目标行为的距离。
  • 在自然语言、国际象棋和路径规划中验证了有效性。
  • 适合用于自动化评估与红队测试,尤其对强模型有潜力。

为帮助评估和理解大语言模型的潜在能力,本文提出使用优化的输入嵌入(即‘软提示’)作为模型与目标行为之间条件距离的度量。该方法旨在作为自动化红队/评估套件的一部分,促进潜在能力的发现,并以可扩展的方式提供关于潜在有害行为可访问性的定量反馈,适用于未来可能具备欺骗性对齐能力的强大模型。实验在自然语言、国际象棋和路径规划任务中展示了基于软提示的评估框架的有效性,并通过广义条件软提示扩展,支持更复杂的任务评估构建。

原文摘要 · Abstract (English)

To help evaluate and understand the latent capabilities of language models, this paper introduces an approach using optimized input embeddings, or 'soft prompts,' as a metric of conditional distance between a model and a target behavior. The technique aims to facilitate latent capability discovery as a part of automated red teaming/evaluation suites and to provide quantitative feedback about the accessibility of potentially concerning behaviors in a way that may scale to powerful future models, including those which may otherwise be capable of deceptive alignment. An evaluation framework using soft prompts is demonstrated in natural language, chess, and pathfinding, and the technique is extended with generalized conditional soft prompts to aid in constructing task evaluations.

模型评估软提示能力探测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。