单一增益比无法区分模型激活迁移中的效力与效能差异。
Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy
- 用跨上下文的EC50差值替代传统增益比,更准确反映迁移效果。
- 五分之四的测试单元中增益比随剂量变化,且存在相反效应方向。
- 适用于需要精准评估模型迁移行为的研究者,尤其关注机制解释时。
在单轮对话中校准添加式激活控制,随后部署至代理框架中。通常报告的指标是增益比 T = Delta_agent / Delta_chat。我们在六种模型上对八类组合进行剂量-响应测试,发现该比值无法识别效力与效能。在自建网格上重建原估计器,其实际范围覆盖了五分之一的可评分单元,且在四分之五的单元中随剂量变动;两个具有相反效力偏移的单元(均在主网格上置信区间清洁)返回的增益区间重叠宽度小于0.08。所有可评分单元在某一剂量下为放大器,另一剂量下则为抑制器,因此需标注测量时的剂量。我们改用位置量度:dEC50 = EC50_agent - EC50_chat,即曲线位置的跨情境差值。该量度在主网格上正负双向显著,跨越四个模型家族(+1.013 [+0.777, +1.273] 对比 -12.368、-10.855、-5.497 和 -886.066),在冻结审计的五个单元中表现优于垂直缩放,在四单元中胜过注册修剪版本。三种通缩解释被测量并拒绝,因符号模式与量级不符。研究同时报告自身实践:一单元被明确隔离,预注册预测器在外部验证中被证伪并公开发布为被驳回;注册补救声明未产生合格单元,报告为无回应;对自身注册清单普查显示,当前代码无法生成已注册分支。最终结论为操作指引而非定理:单一强度下的转移结论无法识别变化本质,因位移与增益不可从单一观测点区分。
原文摘要 · Abstract (English)
Additive activation steering is calibrated in single-turn chat, then deployed inside agent scaffolds. The quantity usually reported for that move is a gain: a ratio of steered effects, T = Delta_agent / Delta_chat. We sweep eight family x arm dose-response cells over six models in both deployment contexts and show this ratio does not identify potency and efficacy. Reconstructing the published estimator in both of its forms on our own grids, its realized range contains 1 in five of five scorable cells, it moves with dose in four of five, and two cells with opposite potency shifts, both CI-clean on the primary grid, return gain intervals overlapping at a width under 0.08. Every scorable cell is an amplifier at one dose and an attenuator at another, so an amplify/attenuate taxonomy reports the dose it was read at. We replace the gain with a location: dEC50 = EC50_agent - EC50_chat, the cross-context difference in a curve location. It is signed both ways on this roster's primary grid, across four model families (+1.013 [+0.777, +1.273] against -12.368, -10.855, -5.497 and -886.066 elsewhere) and beats a vertical rescaling at equal complexity in all five cells of a frozen audit (four under the registered trim). Three deflationary accounts are measured and rejected on sign pattern and magnitude. We report the discipline at the same volume as the result: one cell is quarantined loudly, our pre-registered forecaster was refuted out of sample and is published as refuted, a registered salvage claim produced no qualifying cell and is reported unanswered, and a census of our own register reports registered branches no code here could have emitted. The consequence is a measurement instruction rather than a theorem: a transfer conclusion read at one strength does not identify what changed, because a displacement and a gain are not distinguishable from a single operating point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。