arXiv:2606.31168cs.CRcs.LG2026-06被引 1

三种案例揭示文本型微调模型的记忆检测误判问题

Probe Choice Changes Canary-Memorization Verdicts: Three Post-Hoc Disagreement Case Studies in a Text-Dominant LoRA-Tuned Autoregressive Testbed

论文配图:Probe Choice Changes Canary-Memorization Verdicts: Three Post-Hoc Disagreement Case Studies in a Text-Dominant LoRA-Tuned Autoregressive Testbed
图 1 · 摘自论文原文
  • 用窗口均值负对数似然探针检测记忆,但结果与完整跨度不一致
  • 出现假阳性、假阴性及窗口内下降等三类矛盾情况
  • 建议报告全跨度秘密似然与行为召回率,避免误判

我们在一个基于Qwen2.5-VL-7B的可控小妖测试平台中,审计了一个固定前缀窗口均值负对数似然(K=20)的记忆探测器。报告了三个事后分析案例,其探测结果与全跨度秘密负对数似然或贪婪精确召回存在分歧:案例C3(假阴性,窗口截断)显示,损伤落在窗口外的十六进制令牌上,探针保持平坦而命中率@1下降;案例C4(假阳性,非秘密漂移)显示,探针移动但约99%位于非秘密前缀,秘密跨度和命中率@1未变;案例C5(窗口内模糊下降)显示,探针在低训练基线处下降,而全跨度十六进制仍为正且命中率@1=0。建议:报告(i)全跨度秘密负对数似然,(ii)局部跨度分解,(iii)k≥4时的行为精确召回,以及(iv)干扰探针,方可断言秘密特异性。证据来自单一主干模型的可控小妖;量级具测试平台特异性。

原文摘要 · Abstract (English)

We audit a fixed prefix-window mean-NLL memorization probe (K=20) on a Qwen2.5-VL-7B canary testbed and report three post-hoc cases where it disagrees with full-span secret NLL or greedy exact-recall. C3 (false negative, window truncation): damage lands on hex tokens outside K=20; the probe stays flat while hit@1 drops. C4 (false positive, non-secret drift): the probe moves, but approximately 99% sits on non-secret preamble; the secret span and hit@1 are unchanged. C5 (ambiguous in-window drop): the probe falls on an undertrained baseline while full-span hex is positive and hit@1=0. Recommendation: report (i) full-span secret NLL, (ii) a span-localised decomposition, (iii) behavioural exact-recall at k>=4, and (iv) decoy probes before asserting secret-specificity. Evidence is on controlled canaries in one backbone; magnitudes are testbed-specific.

记忆检测微调模型探针评估小妖测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。