研究离线数据对语言模型探测器泛化能力的影响,发现意图类行为探测易失效。
The Impact of Off-Policy Training Data on Probe Generalisation
- 用合成或离线数据训练探测器,评估其在不同行为上的泛化表现。
- 意图类行为(如策略性欺骗)的探测泛化失败最严重,文本内容类则较稳定。
- 可借助激励数据测试预测真实场景下的探测效果,适合安全监控研究者。
探测技术已成为监控大语言模型(LLMs)的有力手段,可在推理时低成本检测潜在风险行为。然而,许多行为的真实示例稀少,研究人员不得不依赖合成或离线生成的LLM响应来训练探测器。我们系统评估了离线数据对八种不同LLM行为探测泛化性能的影响。通过在多个LLM上测试线性与注意力探测器,发现数据生成策略显著影响探测表现,且差异因行为而异。最大泛化失败出现在以响应‘意图’定义的行为(如策略性欺骗)上,而非仅基于文本内容(如列表使用)。我们提出一种有效测试方法:若探测器能在被诱导生成的数据上表现良好,则其在真实在线数据上也更可能表现优异。基于此,预测当前欺骗探测器在真实监控场景中可能失效。有趣的是,来自差异较大设置的离线数据有时能产生比同源数据更可靠的探测器。这凸显了需发展能应对各类分布偏移的更优监控方法。
原文摘要 · Abstract (English)
Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning behaviours. However, natural examples of many behaviours are rare, forcing researchers to rely on synthetic or off-policy LLM responses for training probes. We systematically evaluate how off-policy data influences probe generalisation across eight distinct LLM behaviours. Testing linear and attention probes across multiple LLMs, we find that training data generation strategy can significantly affect probe performance, though the magnitude varies greatly by behaviour. The largest generalisation failures arise for behaviours defined by response ``intent'' (e.g., strategic deception) rather than text-level content (e.g., usage of lists). We then propose a useful test for predicting generalisation failures in cases where on-policy test data is unavailable: successful generalisation to incentivised data (where the model was coerced) strongly correlates with high performance against on-policy examples. Based on these results, we predict that current deception probes may fail to generalise to real monitoring scenarios. We find that off-policy data can yield more reliable probes than on-policy data from a sufficiently different setting. This underscores the need for better monitoring methods that handle all types of distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。