arXiv:2605.08558cs.LG2026-05

让低精度模拟器自我进化,动态降低高精度评估成本。

Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies

论文配图:Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies
图 1 · 摘自论文原文
  • 设计自适应延续机制,让低精度代理随使用次数提升可靠性。
  • 在合成任务与大模型评价值任务中,显著降低单位成本的误差。
  • 适合需要频繁评估且预算有限的强化学习或模型调优场景。

多保真度多臂赌博机(MF-MAB)允许用不同成本与精度的反馈源评估各选项。传统模型假设保真度差异固定,而现代代理如基于学习的模拟器和大语言模型可通过校准持续改进。本文研究可进化的代理下的自适应多保真度赌博机,聚焦双保真度情形:低保真度源随使用次数变得更可靠。提出选择平均偏差界,将动态低保真观察转化为对高保真目标的改进感知置信区间。设计阈值驱动的自适应延续同伴(TACC)算法,通过有界延续规则决定何时继续低保真采样,何时升级。理论证明实例相关后悔上界:对中等性能臂,自适应延续以有界低保真采样替代对数级高保真确认。在合成赌博机和以大模型为裁判的策略评估任务中验证,延续策略能有效改善成本加权后悔表现。

原文摘要 · Abstract (English)

As an extension of the classical multi-armed bandit problem, multi-fidelity multi-armed bandits (MF-MAB) enable individual arms to be evaluated using diverse feedback sources that vary in both cost and accuracy. Prior stochastic models typically assume fixed low-to-high fidelity discrepancies, whereas modern proxy sources, such as learning-based simulators and Large Language Models (LLMs), can be improved using additional calibration. We investigate adaptive MF-MAB with improving proxy sources, and focus on the canonical two-fidelity case in which the low-fidelity source becomes more informative with repeated use. To capture this dynamic, we introduce a selected-average mismatch bound that converts dynamic low-fidelity observations into improvement-aware confidence bounds for the high-fidelity target. We propose the Threshold-Based Adaptive Continuation Companion (TACC), an optimistic algorithm that uses a bounded continuation rule to decide when low-fidelity sampling remains cost-effective and when to escalate. We prove an instance-dependent regret bound showing that, for detected intermediate arms, adaptive continuation replaces logarithmic high-fidelity confirmation with bounded low-fidelity continuation. Experiments on synthetic bandits and an LLM-as-a-judge policy-evaluation task examine when continuation improves cost-weighted regret.

多保真度强化学习大模型评估自适应决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。