arXiv:2606.08791econ.EMcs.AI2026-06被引 1

通过输入输出分析算法决策者,用协方差分解精确评估其长期损失。

Evaluating AI Investment Strategies

  • 用每期成本与决策的协方差之和,精确计算动态策略的累积后悔值。
  • 在独立同分布成本下,该分解完全成立;非平稳情况下有闭式偏差修正。
  • 适用于平台机制、拍卖和投资策略等需外部审计的序列决策系统。

我们研究仅从可观测输入输出审计黑箱算法决策者的难题。核心结果是:在严格限定条件下,动态策略的累积后悔值可精确分解为各期成本向量与策略决策之间的协方差之和。该结果将Aldridge(2026)的单周期恒等式拓展至随机动态规划的多周期设定。在独立同分布成本和均值无偏马尔可夫策略下,该恒等式严格成立;针对非平稳和时变情形,推导出闭式偏差校正;并建立折扣时间窗情形下的对应关系。协方差后悔函数满足贝尔曼递归,与标准强化学习算法兼容;滚动窗口策略的估计误差偏差为O(d/w)。该分解对战略环境中的算法审计有直接意义:在平台机制设计中,无需访问代理人私有类型即可提供基于福利的审计指标;在重复博弈中,协方差降低是策略改进的充分条件;在采购与广告拍卖中,偏差校正量化了策略虚报导致的福利损失。所提出的轨迹估计器具有一致性,渐近正态且具有异质自相关稳健方差,计算复杂度为O(T·nd),使该方法成为平台机制、算法投资策略及任何受外部绩效审查的序列决策系统的可行、模型无关审计工具。

原文摘要 · Abstract (English)

We study the problem of auditing a black-box algorithmic decision-maker from observable inputs and outputs alone. Our main result is an exact decomposition: under precisely characterized conditions, the cumulative \emph{regret} of a dynamic policy equals the sum of per-period covariances between the cost vector and the policy's decision. This extends the single-period identity of Aldridge~(2026) to the full multi-period setting of stochastic dynamic programming. We prove the identity holds exactly under i.i.d. costs and mean-unbiased Markov policies, derive closed-form bias corrections for non-stationary and time-varying cases, and establish the discounted-horizon analog. A Bellman recursion for the covariance regret functional connects the result to standard reinforcement learning algorithms; for rolling-window policies, the estimation-error bias is $O(d/w)$. The decomposition has direct implications for algorithmic auditing in strategic environments: in platform mechanism design, it provides a welfare-based audit metric without access to the agent's private type; in repeated games, covariance reduction is a sufficient condition for policy improvement; in procurement and ad auctions, the bias correction quantifies welfare loss from strategic misreporting. The associated trajectory estimator is consistent, asymptotically normal with HAC variance, and computable in $O(T \cdot nd)$ time. This makes the proposed approach a tractable, model-free audit tool for platform mechanisms, algorithmic portfolio strategies, and any sequential decision system subject to external performance review.

算法审计后悔分析决策优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。