提出可预测选择性干预路径的内存定位方法,提升模型控制精度与安全性。
Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

- 基于内部信号预测干预路径,区分语义邻近与能力损伤。
- 在第7层达到13.1%目标命中率,优于随机基线3.6个百分点。
- 适用于需要低剂量、高风险可控干预的场景,如医疗或金融决策。
激活调控将局部表征转化为控制方向,但定位本身无法揭示方向是否具有选择性操作区间。本文提出预测性内存定位(PML),将测量网格干预路径作为内存定位的预测目标。PML分离随机校准的目标移动与语义邻近及能力损伤,并通过强度不重叠的低剂量因果响应,对比静态定位与监督几何。冻结研究涵盖九个数据集、十四类领域共3000条记录,生成3万条独立记录-方向-层路径及21万次路径强度评估。在第7层,几何推导的RFM/AGOP方向在记录配对自助法下实现13.1%目标任意命中率和12.3%干净任意命中率,分别优于随机基线3.6和3.4个百分点。在记录、数据集、领域分组分割下,|α|=0.1时的响应是|α|∈{0.25,0.5}离散强度结果最强信号。在保留记录上,基于预测器的系数选择器优于训练调优的固定强度策略,提升效用并减少语义邻近损伤,且在密集扫描中避免多数评估。在三个残差归一化匹配的基线模型中,学习方向保持选择性路径优势,低剂量响应获得0.801–0.828的记录保留宏AUROC。因此,PML将内存定位转化为可验证的边际选择性结果预测与风险感知干预决策。
原文摘要 · Abstract (English)
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record-direction-layer paths and 210,000 distinct path-strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at $|α|=0.1$ are the strongest signal for outcomes at disjoint strengths $|α|\in\{0.25,0.5\}$. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801-0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。