研究顶K接口对模型分布提取的限制,发现能力提取仍可能成功。
Identified-Set Geometry of Distributional Model Extraction under Top-$K$ Censored API Access
- 基于顶K输出推断教师模型分布的可识别集几何结构。
- 全量日志提取恢复56%私有能力,生成式方法达96%。
- 接口限制虽阻碍分布还原,但不影响能力窃取,适合安全评估者阅读。
现代大模型API通常仅返回顶-K logits并屏蔽其余词汇。本文研究在此访问模式下,逐位置分布恢复的极限。对于截断阈值τ,相容的教师分布构成一个可识别集,其总变差直径精确为U_K=(V-K)exp(τ)/(Z_A+(V-K)exp(τ)),其中Z_A为观测到的分区函数。针对KL距离恢复,给出了可计算的二元端点下界与渐近匹配的小模糊上界,并扩展至参考感知攻击场景。在Qwen3数学推理教师模型上的实验显示分层提取层级:任务相关顶-K蒸馏恢复12%私有能力,全日志蒸馏在99% KL闭合下恢复56%,生成式提取恢复96%。因此,顶-K截断虽限制分布恢复,但不阻止能力提取,揭示了提示仅日志蒸馏中保真度与迁移性的分离。
原文摘要 · Abstract (English)
Modern LLM APIs often reveal only top-$K$ logit scores and censor the remaining vocabulary. We study the per-position distribution-recovery limits of this access model. For censoring threshold $τ$, the compatible teacher distributions form an identified set whose total-variation diameter is exactly $U_K=(V-K)\exp(τ)/(Z_A+(V-K)\exp(τ))$, where $Z_A$ is the observed partition function. For KL recovery, we give a computable binary-endpoint lower bound and an asymptotically matching small-ambiguity upper bound, with an extension to reference-aware attackers. Experiments on a Qwen3 math-reasoning teacher reveal a layered extraction hierarchy: on-task top-$K$ distillation recovers 12% of private capability, full-logit distillation recovers 56% despite 99% KL closure, and generation-based extraction recovers 96%. Top-$K$ censoring therefore limits per-position distribution recovery but does not by itself prevent capability extraction, separating fidelity from transfer in prompt-only logit distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。