用影响样本重写回复,比重调权重更能有效改变大模型行为。
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

- 用影响函数选样后,替换响应而非调整权重
- 重写使模型行为改变更强更持久,且可正可反
- 适合研究模型可解释性与安全可控性的研究人员
训练数据溯源(TDA)旨在识别影响模型行为的训练样本,但其干预效果取决于选择哪些样本以及如何修改。影响函数(IF)估计微小权重调整下的行为变化,但使用IF选出的样本在传统权重干预下常不如随机选择有效。这引发疑问:是这些高影响力样本本身缺乏干预价值,还是权重调整无法发挥其潜力?本文提出影响引导的响应重写方法:利用IF识别干预目标,固定提示语,仅替换其响应为一致或对立的监督信号。在四个开源大模型上,以知识性回避为测试基准,对比相同影响选样下的重写与重调。结果显示,重写产生更强、更持久、双向的行为改变,而重调效果微弱且不一致。进一步分析表明,影响选样比其他方法更具重写杠杆效应,且变化集中于目标相关行为。该结论在安全拒答任务中同样成立。结果揭示了影响估计所捕捉的局部权重效应与样本实际具备的广泛干预潜力之间的差异,呼吁对TDA方法进行干预感知评估。
原文摘要 · Abstract (English)
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。