arXiv:2509.00095cs.LGcs.NE2025-09

用强化学习优化企业预算分配,兼顾真实支出习惯与动态变化。

Financial Decision Making using Reinforcement Learning with Dirichlet Priors and Quantum-Inspired Genetic Optimization

  • 引入狄利克雷分布模拟财务环境变化,增强模型随机性。
  • 量子启发遗传算法提升性能,未见数据上相似度达0.9990。
  • 适合关注智能财务决策与算法优化的研究者和从业者。

传统预算分配模型难以应对现实金融数据的随机性和非线性特征。本研究提出一种融合狄利克雷先验与量子启发遗传优化的混合强化学习框架,用于动态预算分配。基于苹果公司2009至2025年季度财务数据,该智能体学习在研发与销售、管理费用间分配预算,以最大化盈利并遵循历史支出模式,同时引入L2惩罚项避免不切实际的偏离。状态演化采用狄利克雷分布以模拟不断变化的财务背景。为跳出局部最优并提升泛化能力,训练后的策略通过参数化量子比特旋转电路实现量子变异的遗传算法进行精炼。每代的奖励与惩罚被记录以可视化收敛过程与策略行为。在未见的财政数据上,模型与实际分配高度一致(余弦相似度0.9990,KL散度0.0023),证明深度强化学习、随机建模与量子启发启发式方法结合在自适应企业预算管理中的潜力。

原文摘要 · Abstract (English)

Traditional budget allocation models struggle with the stochastic and nonlinear nature of real-world financial data. This study proposes a hybrid reinforcement learning (RL) framework for dynamic budget allocation, enhanced with Dirichlet-inspired stochasticity and quantum mutation-based genetic optimization. Using Apple Inc. quarterly financial data (2009 to 2025), the RL agent learns to allocate budgets between Research and Development and Selling, General and Administrative to maximize profitability while adhering to historical spending patterns, with L2 penalties discouraging unrealistic deviations. A Dirichlet distribution governs state evolution to simulate shifting financial contexts. To escape local minima and improve generalization, the trained policy is refined using genetic algorithms with quantum mutation via parameterized qubit rotation circuits. Generation-wise rewards and penalties are logged to visualize convergence and policy behavior. On unseen fiscal data, the model achieves high alignment with actual allocations (cosine similarity 0.9990, KL divergence 0.0023), demonstrating the promise of combining deep RL, stochastic modeling, and quantum-inspired heuristics for adaptive enterprise budgeting.

强化学习预算分配量子启发金融建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。