用知识边界指纹技术,低成本检测大模型接口是否被替换。
KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

- 通过模型在知识边界附近的稳定数值召回特征进行指纹识别。
- 16个生产环境接口中,准确识别出155次经济驱动的模型替换,零误判同源控制。
- 适合关注模型接口安全、防止服务冒充的研究者与企业用户。
中介型大语言模型(LLM)接口日益普遍,但用户无法验证声称的接口是否真正提供指定模型。我们提出KBF,一种低成本黑盒审计协议,利用模型在知识边界附近稳定的数值召回特征作为指纹。在16个生产级LLM接口上,KBF成功识别全部155次由经济动机引发的模型替换,且未误判任何同源控制;在部署变化下保持稳定;当仅5-10%流量被替换时即可检测到高分离混合路由攻击;在六平台影子接口审计中,27个模型单元中有7个在统计上与参考端点不一致,异常集中于高端Claude接口。
原文摘要 · Abstract (English)
Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. Across 16 production LLM endpoints, KBF flags all 155 economically relevant substitutions without rejecting any same-model controls, remains stable under deployment variation, detects high-separation mixed-routing attacks when only 5-10% of traffic is substituted, and finds that 7 of 27 platform model cells in a six-platform shadow API audit are statistically inconsistent with their reference endpoints, with inconsistencies concentrated on premium Claude endpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。