提出多轮对话中激发大模型特定行为的新方法,提升评估效率与发现能力。
Eliciting Behaviors in Multi-Turn Conversations
- 构建三类交互方式的分析框架,统一单轮与多轮激发机制。
- 在线方法仅用数千次交互即达成45%/19%/77%的成功率,显著优于静态基准。
- 适合关注大模型行为评估与动态测试基准的研究者使用。
在对话场景中识别大型语言模型(LLMs)的具体甚至复杂行为对评估至关重要。近期研究提出新方法,通过自然语言提示激发目标模型的行为,但主要集中在单轮设置。本文研究多轮对话中的行为激发问题,首先提出一个分析框架,将现有方法按与目标模型的交互方式分为三类:仅依赖先验知识、离线交互和在线学习。接着引入在线方法的广义多轮形式,统一单轮与多轮激发。我们在自动生成多轮测试案例上评估三类方法,分析查询预算(与目标模型的交互次数)与成功率(行为激发输入的发现率)之间的权衡。结果表明,在线方法在三个任务中仅用数千次查询即可达到45%/19%/77%的平均成功率,而现有静态多轮对话基准几乎无法发现失败案例。本工作揭示了行为激发在多轮对话评估中的新应用,并呼吁社区转向动态基准。
原文摘要 · Abstract (English)
Identifying specific and often complex behaviors from large language models (LLMs) in conversational settings is crucial for their evaluation. Recent work proposes novel techniques to find natural language prompts that induce specific behaviors from a target model, yet they are mainly studied in single-turn settings. In this work, we study behavior elicitation in the context of multi-turn conversations. We first offer an analytical framework that categorizes existing methods into three families based on their interactions with the target model: those that use only prior knowledge, those that use offline interactions, and those that learn from online interactions. We then introduce a generalized multi-turn formulation of the online method, unifying single-turn and multi-turn elicitation. We evaluate all three families of methods on automatically generating multi-turn test cases. We investigate the efficiency of these approaches by analyzing the trade-off between the query budget, i.e., the number of interactions with the target model, and the success rate, i.e., the discovery rate of behavior-eliciting inputs. We find that online methods can achieve an average success rate of 45/19/77% with just a few thousand queries over three tasks where static methods from existing multi-turn conversation benchmarks find few or even no failure cases. Our work highlights a novel application of behavior elicitation methods in multi-turn conversation evaluation and the need for the community to move towards dynamic benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。