根据可验证输出动态调整大模型推理成本,平衡质量与开销。
Belief-Guided Inference Control for Large Language Model Services via Verifiable Observations

- 构建输入输出的可信观测通道,生成响应可靠性信念状态。
- 在多任务测试中实现更优的质量-成本权衡与风险校准。
- 适合对响应可靠性要求高的生产级大模型服务场景。
在黑盒大语言模型服务中,响应可靠性仅在决策时刻部分可观测,而更强的推理路径会带来高昂计算成本,形成带预算的序列决策问题:每个请求需判断默认低耗响应是否足够可靠,或是否应追加计算以提升质量。本文提出Verifiable Observations for Risk-aware Inference Control(Veroic)框架,将请求时控制建模为部分可观测马尔可夫决策过程,以捕捉部分可观测性与预算的序列耦合。该框架通过聚合异构质量信号,从输入输出对构建轻量级可验证观测通道,生成对隐含响应可靠性的信念状态,并由预算感知策略决定是否返回默认输出或触发高成本推理路径。在多种任务上的实验表明,Veroic在质量-成本权衡、风险估计与校准、长周期推理控制方面均优于现有基线。
原文摘要 · Abstract (English)
In black-box large language model (LLM) services, response reliability is often only partially observable at decision time, while stronger inference pathways incur substantial computational cost, inducing a budgeted sequential decision problem: for each request, the system should decide whether the default low-cost response is sufficiently reliable or whether additional computation should be allocated to improve response quality. In this paper, we propose \textbf{Ver}ifiable \textbf{O}bservations for Risk-aware \textbf{I}nference \textbf{C}ontrol (\textsc{Veroic}), a framework for adaptive inference control in black-box LLM settings, which formulates request-time control as a \textit{partially observable Markov decision process} to capture partial observability and sequential budget coupling. It constructs a lightweight verifiable observation channel from the input-output pair by aggregating heterogeneous quality signals into a belief state over latent response reliability, which is then used by a budget-aware policy to decide whether to return the default output or trigger a higher-cost inference pathway. Experiments on diverse tasks show that \textsc{Veroic} achieves improved quality-cost trade-offs, stronger risk estimation and calibration, and more robust long-horizon inference control than competitive baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。