提出反事实似然测试,精准区分私有推理中的间接影响。
Counterfactual Likelihood Tests for Indirect Influence in Private Reasoning Channels
- 用同长替换块构造反事实场景,测量下游输出变化。
- 发现私有通道间存在单向影响,且仅通过公开通信路径传递。
- 适合关注隐私推理安全性的研究人员使用。
推理系统日益将中间计算分为私有与公开通道,导致评估案例在文本记录上相似:独立共推导、直接访问私有内容,以及通过公开沟通的间接影响。本文提出一种反事实似然测试,用于测量私有推理通道间的影响力。方法将上游私有模块替换为长度匹配的供体块,固定公开标记序列和下游目标,测量下游目标负对数似然的变化。在7B角色-通道推理模型上的验证显示,文本探针不可靠:原始n元语法重叠高估泄露,修正后重叠仍噪声大,而金丝雀复现报告无区分能力。反事实似然能有效分离未遮蔽与遮蔽条件,长度匹配控制了RoPE位置混淆。在强化遮蔽验证中,反向B→A影响接近零,而A→B影响仍通过公开语音隐藏状态传递。跨三个检查点、五个随机种子、13,734个有效方向对比的多检查点验证重现此不对称性。图分离控制(阻断私有→公开传输边)在所有13,734次控制评估中产生比特级相同的自然与反事实得分,确认所测公共通道路径是反事实信号的完整载体。结果表明,私有通道评估应分别报告直接与间接影响,反事实似然探针可作为测量边界的标准方法。
原文摘要 · Abstract (English)
Reasoning systems increasingly separate intermediate computation into private and public channels, creating evaluation cases that look similar in transcripts: independent co-derivation, direct access to private content, and indirect influence through public communication. This paper presents a counterfactual likelihood test for measuring influence between private reasoning channels. The method replaces an upstream private block with a length-matched donor block, holds the public token sequence and downstream target fixed, and measures the downstream target's negative-log-likelihood shift. On a 7B role-channel reasoning model used for validation, textual probes are unreliable: raw n-gram overlap overstates leakage, corrected overlap remains noisy, and canary reproduction reports no discrimination. Counterfactual likelihood separates unmasked and masked conditions, while length matching controls a RoPE positional confound. In the hardened masked validation, reverse B-to-A influence is near zero, while A-to-B influence persists through public-speech hidden states. A multi-checkpoint validation across three checkpoints, five seeds, and 13,734 valid directional contrasts replicates this asymmetry. A graph-separation control that blocks private-to-public carrier edges produces bit-identical natural and counterfactual scores across all 13,734 control evaluations, identifying the tested public-channel pathway as the complete carrier of the measured counterfactual signal under the implemented role-visibility mask. The results show that private-channel evaluation should report direct and indirect influence separately, and that counterfactual likelihood probes provide a practical default for measuring these boundaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。