arXiv:2410.17245cs.AIcs.CL2024-10中稿 · NeurIPS被引 36

提出四条评估大模型行为干预的新标准,让效果评测更客观可靠。

Towards Reliable Evaluation of Behavior Steering Interventions in LLMs

  • 基于下游任务相似上下文和模型似然值设计量化评估流程
  • 发现部分行为干预效果比之前报告的差,验证了新方法的敏感性
  • 适合关注模型可控性与评估方法论的研究者

表示工程方法在高效引导大模型行为方面展现出潜力,但现有评估多依赖主观演示,缺乏定量客观指标。本文提出四个当前评估缺失的关键属性:(i) 使用与下游任务足够相似的上下文评估干预质量;(ii) 考虑模型似然值;(iii) 支持不同目标行为间的标准化比较;(iv) 提供基线对比。我们据此构建了一套评估流程,提供定量与可视化分析,用于评估两种表示工程方法在提升真实性与可纠正性等行为上的有效性,结果表明部分干预的实际效果低于先前报告水平。

原文摘要 · Abstract (English)

Representation engineering methods have recently shown promise for enabling efficient steering of model behavior. However, evaluation pipelines for these methods have primarily relied on subjective demonstrations, instead of quantitative, objective metrics. We aim to take a step towards addressing this issue by advocating for four properties missing from current evaluations: (i) contexts sufficiently similar to downstream tasks should be used for assessing intervention quality; (ii) model likelihoods should be accounted for; (iii) evaluations should allow for standardized comparisons across different target behaviors; and (iv) baseline comparisons should be offered. We introduce an evaluation pipeline grounded in these criteria, offering both a quantitative and visual analysis of how effectively a given method works. We use this pipeline to evaluate two representation engineering methods on how effectively they can steer behaviors such as truthfulness and corrigibility, finding that some interventions are less effective than previously reported.

行为控制模型评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。