揭露影子API虚假承诺,挑战学术研究可信度
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
- 对比官方与影子API输出,发现性能差距达47.21%
- 45.83%指纹测试中影子API无法通过身份验证
- 适合关注AI研究可复现性与数据真实性的学者
前沿大语言模型(如GPT-5、Gemini-2.5)因定价高、支付门槛和区域限制难以获取,催生了大量声称绕过这些限制的第三方‘影子API’。尽管其被187篇学术论文使用,但其输出是否真实可靠仍不明。本文首次系统审计官方API与对应影子API。识别出17个影子API,其中最流行者获5,966次引用、58,639个GitHub星标(截至2025年12月6日)。通过功能、安全与模型验证多维度检测,发现性能偏差最高达47.21%,安全行为高度不可预测,45.83%指纹测试中身份验证失败。此类欺骗行为严重威胁研究可复现性,损害用户利益并损害官方模型声誉。
原文摘要 · Abstract (English)
Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions. These limitations drive the proliferation of $\textit{shadow APIs}$, third-party services that claim to provide access to official model services without regional limitations via indirect access. Despite their widespread use, it remains unclear whether shadow APIs deliver outputs consistent with those of the official APIs, raising concerns about the reliability of downstream applications and the validity of research findings that depend on them. In this paper, we present the first systematic audit between official LLM APIs and corresponding shadow APIs. We first identify 17 shadow APIs that have been utilized in 187 academic papers, with the most popular one reaching 5,966 citations and 58,639 GitHub stars by December 6, 2025. Through multidimensional auditing of three representative shadow APIs across utility, safety, and model verification, we uncover both indirect and direct evidence of deception practices in shadow APIs. Specifically, we reveal performance divergence reaching up to $47.21\%$, significant unpredictability in safety behaviors, and identity verification failures in $45.83\%$ of fingerprint tests. These deceptive practices critically undermine the reproducibility and validity of scientific research, harm the interests of shadow API users, and damage the reputation of official model providers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。