arXiv:2511.01023eess.SPcs.AI2025-11被引 4

模型初始化种子决定隐性信息泄露程度,子空间对齐是关键

Seed-Induced Uniqueness in Transformer Models: Subspace Alignment Governs Subliminal Transfer

  • 通过子空间对齐分析隐性迁移机制
  • 相同种子下泄露率τ≈0.24,不同种子仅τ≈0.12-0.13
  • 适合关注模型安全与隐私保护的研究者

我们分析了Transformer模型中的隐性迁移现象,即教师模型将隐藏特征以线性方式传递给学生模型,且不损害主任务性能。以往研究多依赖全局表示相似性(如中心核对齐CKA)解释可迁移性。使用包含解耦公共与私有标签的合成语料库,在匹配与独立随机初始化下进行学生模型蒸馏。结果发现,迁移强度取决于特征区分性子空间内的对齐:同种子学生继承该对齐,泄露率τ≈0.24;而不同种子学生尽管全局CKA > 0.9,但额外准确率显著降低,τ≈0.12–0.13。通过子空间级CKA诊断与残差探针验证,证明泄露行为与特征区分性子空间对齐相关,而非全局表示相似性。安全控制手段(投影惩罚、对抗反转、错误理由正则化)在不损害公共任务保真度的前提下,有效降低同源模型的泄露。结果确立了种子诱导的独特性为一种鲁棒性属性,并主张在多模型部署中采用子空间感知诊断。

原文摘要 · Abstract (English)

We analyze subliminal transfer in Transformer models, where a teacher embeds hidden traits that can be linearly decoded by a student without degrading main-task performance. Prior work often attributes transferability to global representational similarity, typically quantified with Centered Kernel Alignment (CKA). Using synthetic corpora with disentangled public and private labels, we distill students under matched and independent random initializations. We find that transfer strength hinges on alignment within a trait-discriminative subspace: same-seed students inherit this alignment and show higher leakage {τ\approx} 0.24, whereas different-seed students -- despite global CKA > 0.9 -- exhibit substantially reduced excess accuracy {τ\approx} 0.12 - 0.13. We formalize this with subspace-level CKA diagnostic and residualized probes, showing that leakage tracks alignment within the trait-discriminative subspace rather than global representational similarity. Security controls (projection penalty, adversarial reversal, right-for-the-wrong-reasons regularization) reduce leakage in same-base models without impairing public-task fidelity. These results establish seed-induced uniqueness as a resilience property and argue for subspace-aware diagnostics for secure multi-model deployments.

模型安全隐性迁移子空间对齐隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。