用数学方法检测大模型服务是否谎报调用次数,确保用户不被多收费。
Auditing Pay-Per-Token in Large Language Models
- 基于鞅理论设计审计框架,持续查询验证输出令牌数
- 少于70次输出即可发现欺诈行为,误判率低于5%
- 适用于所有主流大模型,适合关心计费公平的用户
数百万用户依赖云端服务获取最先进的大语言模型。然而,近期研究表明,当前主流的按令牌计费机制会激励服务商策略性地伪造或虚报生成结果所使用的令牌数。本文提出一种基于鞅理论的审计框架,使可信第三方审计者通过连续查询服务提供商,可无条件检测令牌虚报行为,且几乎不会将诚实的提供者误判为不诚实。我们使用来自Llama、Gemma和Ministral系列的多个大模型,结合一个流行的众包评测平台上的输入提示,在多种(不)诚实报告策略下进行实验。结果表明,该框架在观察少于约70个输出后即可检测到不诚实的服务商,同时将误判忠实服务商的概率控制在α=0.05以下。
原文摘要 · Abstract (English)
Millions of users rely on a market of cloud-based services to obtain access to state-of-the-art large language models. However, it has been very recently shown that the de facto pay-per-token pricing mechanism used by providers creates a financial incentive for them to strategize and misreport the (number of) tokens a model used to generate an output. In this paper, we develop an auditing framework based on martingale theory that enables a trusted third-party auditor who sequentially queries a provider to detect token misreporting. Crucially, we show that our framework is guaranteed to always detect token misreporting, regardless of the provider's (mis-)reporting policy, and not falsely flag a faithful provider as unfaithful with high probability. To validate our auditing framework, we conduct experiments across a wide range of (mis-)reporting policies using several large language models from the $\texttt{Llama}$, $\texttt{Gemma}$ and $\texttt{Ministral}$ families, and input prompts from a popular crowdsourced benchmarking platform. The results show that our framework detects an unfaithful provider after observing fewer than $\sim 70$ reported outputs, while maintaining the probability of falsely flagging a faithful provider below $α= 0.05$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。