arXiv:2606.17383q-fin.RMcs.AI2026-06

为自主AI系统构建基于部分可观测马尔可夫决策过程的验证框架。

Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation

  • 用POMDP分解智能体决策流程,分步验证信念、预测与策略
  • 实证显示隐状态推断显著提升决策质量,且结果对参数变化鲁棒
  • 适合金融风控、自动驾驶等高风险自主系统验证场景

自主人工智能系统引入新型模型风险。与传统预测模型不同,智能体持续获取信息,形成对环境隐状态的信念,生成预测,选择行动并动态调整行为。现有验证方法主要关注预测准确率,难以评估决策过程质量。本文提出基于部分可观测马尔可夫决策过程(POMDP)的智能体模型验证框架,将自主决策分解为信息、信念、预测、行动与效用五个环节,支持独立验证。大语言模型被形式化为近似贝叶斯滤波算子,并建立涵盖状态空间、滤波、预测、策略、效用设定与参数风险的模型风险分类体系。通过投资组合管理案例研究验证:智能体从市场与宏观经济信息中推断隐市场状态,生成条件信念预测,并基于Black--Litterman框架构建投资组合。实证包含性能分析、信念校准诊断、覆盖率检验、消融实验与参数敏感性分析。结果显示,隐状态推断独立贡献于决策质量,且核心结论在广泛参数范围内保持稳健。论文主要贡献在于将成熟的模型风险管理概念扩展至自主AI系统,提供可操作的验证、治理与监控基础。

原文摘要 · Abstract (English)

Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents continuously acquire information, form beliefs regarding latent states of the environment, generate forecasts, select actions, and adapt their behavior over time. Existing validation methodologies focus primarily on predictive accuracy and therefore provide limited insight into the quality of the underlying decision process. This paper proposes a model validation framework for agentic AI based on Partially Observable Markov Decision Processes (POMDPs). The framework decomposes autonomous decision making into information, beliefs, forecasts, actions, and utility, allowing each component to be validated independently. Large language models (LLMs) are formalized as approximate Bayesian filtering operators, and a model-risk taxonomy is developed encompassing state-space, filtering, forecast, policy, utility-specification, and parameter risks. The model risk validation methodology is demonstrated through a portfolio-management case study in which an agent infers latent market regimes from market and macroeconomic information, generates belief-conditioned forecasts, and constructs portfolios using a Black--Litterman framework. Empirical validation combines performance analysis, belief calibration diagnostics, coverage tests, ablation studies, and parameter-sensitivity analysis. The results indicate that latent-state inference contributes independently to decision quality and that the principal conclusions remain robust across a broad range of parameter values. The principal contribution of the paper is a practical framework for extending established model risk management concepts to autonomous AI systems and providing a rigorous foundation for their validation, governance, and monitoring.

自主AIPOMDP模型验证风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。