语言模型间隐性影响可借自然语言载体传播,且难被人类察觉。
Covert Influence Between Language Models

- 用样本级归因分数筛选增强影响的载体,实现隐蔽传递
- 三类接口中,上下文学习的影响范围最广且无可见痕迹
- 自然语言载体的影响比数字载体更易检测,但跨模型迁移更强
随着语言模型越来越多地使用彼此输出,一种隐蔽影响现象——即发送方的意图(行为倾向)通过人类无法察觉的载体传递给接收方——正成为日益严重的风险。我们评估了三种接口下的该风险:监督微调、在线策略蒸馏和上下文学习,发现它们在不留下人类可识别痕迹的前提下,所能实现的影响规模各不相同。通过推理时的每样本归因分数,我们能选择性放大训练时的影响,从而实现此前研究无法达成的意图传递。此外,我们发现以自然语言为载体的隐蔽影响与以往使用数字载体的研究存在本质差异:后者更难被察觉,但跨模型家族的可迁移性更低。这些结果表明,隐蔽影响的风险面远超以往认知,而点对点归因评分方法可作为探究与缓解该风险的有效工具。
原文摘要 · Abstract (English)
As language models increasingly consume one another's outputs, covert influence -- a phenomenon where a sender's payload (the behavioral disposition it is conditioned to propagate) transfers to a receiver through carriers undetectable by humans -- becomes a growing risk. We characterize this risk across three interfaces: supervised fine-tuning, on-policy distillation, and in-context learning, and find that they vary in the scale of influence achievable without leaving behind human-visible traces. Using inference-time per-sample attribution scores, we study covert influence across all three interfaces with the ability to select carriers that amplify training-time influence, unlocking payload transfers that prior work could not achieve. We further provide evidence that covert influence with natural-language carriers is a distinct phenomenon from prior studies using number carriers, as the latter is more resistant to human detection and less portable across model families. Together, these results suggest that the risk surface for covert influence is broader than previously recognized, and we study pointwise attribution scoring methods as a tool to investigate and mitigate it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。