对比两种模型行为引导方法,发现不同任务下各有优劣。
Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
- 用函数向量(底向上)和上下文向量(顶向下)分别引导模型行为。
- 上下文向量在改变行为上更有效,函数向量在精确任务中表现更好。
- 研究为未来模型可解释性评估提供新方向,适合关注模型控制的研究者。
大语言模型可解释性研究的一个关键目标是开发稳健引导模型实现期望行为的方法。为此,提出了两种不同的解释路径——‘底向上’与‘顶向下’,但两者之间的定量比较仍较少。本文通过案例研究,对比了两类代表性向量引导方法:底向上方法函数向量(FV;arXiv:2310.15213)与顶向下方法上下文向量(ICV;arXiv:2311.06668)。尽管二者均旨在捕捉广泛上下文学习任务的紧凑表示,我们发现它们仅在特定类型任务中有效:ICVs 在行为迁移方面优于 FVs,而 FVs 在需要更高精度的任务中表现更佳。该结果对今后引导方法的评估及顶向下与底向上引导机制的进一步研究具有启示意义。
原文摘要 · Abstract (English)
A key objective of interpretability research on large language models (LLMs) is to develop methods for robustly steering models toward desired behaviors. To this end, two distinct approaches to interpretability -- ``bottom-up" and ``top-down" -- have been presented, but there has been little quantitative comparison between them. We present a case study comparing the effectiveness of representative vector steering methods from each branch: function vectors (FV; arXiv:2310.15213), as a bottom-up method, and in-context vectors (ICV; arXiv:2311.06668) as a top-down method. While both aim to capture compact representations of broad in-context learning tasks, we find they are effective only on specific types of tasks: ICVs outperform FVs in behavioral shifting, whereas FVs excel in tasks requiring more precision. We discuss the implications for future evaluations of steering methods and for further research into top-down and bottom-up steering given these findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。