数据中隐藏的潜台词可被提取,影响模型行为
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- 通过日志线性选择法,从数据中挖掘隐藏信号
- 实验证明选子集能让模型产生特定偏好或跨语言响应
- 方法通用,适用于不同模型架构,适合安全研究者
训练现代大语言模型已涉及大量算法与数据集,以诱导特定行为,因此理解数据对模型属性的影响至关重要。近期实验发现,数据能传递无法从单个数据点直接观察到的信号,挑战了以数据为中心的模型训练理解,并暗示此类现象缺乏基础解释。受大语言模型线性结构研究启发,我们揭示了一种通用机制,可使通用数据集中出现隐藏语义。提出日志线性选择(LLS)方法,指导如何从通用偏好数据集中选取子集,以引发广泛隐藏效应。将LLS应用于真实数据集,发现所选子集可使模型表现出特定偏好、对数据外语言作出响应,或扮演新身份。关键的是,该效应在不同架构模型间持续存在,支持其普遍性。
原文摘要 · Abstract (English)
Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understand the effects of datasets on the model's properties. This is exacerbated by recent experiments that show datasets can transmit signals that are not directly observable from individual datapoints, posing a conceptual challenge for dataset-centric understandings of LLM training and suggesting a missing fundamental account of such phenomena. Towards understanding such effects, inspired by recent work on the linear structure of LLMs, we uncover a general mechanism through which hidden subtexts can arise in generic datasets. We introduce Logit-Linear-Selection (LLS), a method that prescribes how to select subsets of a generic preference dataset to elicit a wide range of hidden effects. We apply LLS to discover subsets of real-world datasets so that models trained on them exhibit behaviors ranging from having specific preferences, to responding to prompts in a different language not present in the dataset, to taking on a different persona. Crucially, the effect persists for the selected subset, across models with varying architectures, supporting its generality and universality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。