arXiv:2506.15735cs.AIcs.LG2025-06

让大模型按指定方式响应,同时保持语言自然。

ContextBench: Modifying Contexts for Targeted Latent Activation

  • 通过修改上下文精准触发模型特定行为或特征。
  • 新方法在激发目标特征与语言流畅性上均达顶尖水平。
  • 适合安全检测、可控生成等场景研究者参考。

识别能触发语言模型特定行为或潜在特征的输入,可广泛应用于安全领域。我们研究了一类能够生成目标明确且语言流畅的输入以激活特定潜在特征或引发模型行为的方法。将该方法形式化为上下文修改,并提出 ContextBench——一个评估核心能力与安全应用潜力的基准。评估框架同时衡量激发强度(潜在特征或行为的激活程度)与语言流畅性,揭示当前最先进方法难以平衡二者。我们通过引入 LLM 辅助与扩散模型图像修复技术增强进化提示优化(EPO),验证其在激发效果与流畅性平衡上达到当前最优表现。

原文摘要 · Abstract (English)

Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We formalise this approach as context modification and present ContextBench -- a benchmark with tasks assessing core method capabilities and potential safety applications. Our evaluation framework measures both elicitation strength (activation of latent features or behaviours) and linguistic fluency, highlighting how current state-of-the-art methods struggle to balance these objectives. We enhance Evolutionary Prompt Optimisation (EPO) with LLM-assistance and diffusion model inpainting, and demonstrate that these variants achieve state-of-the-art performance in balancing elicitation effectiveness and fluency.

大模型安全提示优化生成控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。