用元强化学习预训练,快速生成个性化理财方案。
A Meta Reinforcement Learning Approach to Goals-Based Wealth Management

- 在数千个理财目标问题上预训练,实现零样本泛化
- 新问题推理仅需0.02秒,平均达最优效用的97.8%
- 适用于复杂市场变化,能解决动态规划难处理的大状态空间
受零样本元学习与基础模型预训练启发,我们提出一种元强化学习方法(记为MetaRL),在数千个基于目标的财富管理(GBWM)问题上进行预训练。每个GBWM问题涉及多年期投资决策,投资者每年需选择投资组合并决定是否实现当年出现的不同财务目标。目标是最大化实现目标带来的总预期效用。该方法在推理时无需重新训练,可在数百毫秒内为新问题生成近似最优的动态投资组合与目标达成策略,平均达到动态规划确定的最优效用的97.8%。结果对资本市场制度变化具有显著鲁棒性,即使训练时仅使用单一市场状态也表现良好。此外,该方法可解决动态规划因计算成本过高而无法处理的大状态空间问题。
原文摘要 · Abstract (English)
Applying concepts related to zero-shot meta-learning and pre-training of foundation models, we develop a meta reinforcement learning approach (denoted MetaRL) that is pre-trained on thousands of goals-based wealth management (GBWM) problems. Each GBWM problem involves a multiple year scenario over which the investor looks to optimally choose an investment portfolio each year and choose to fulfill all, some, or none of the different financial goals that arise each year. These choices seek to maximize the expected total investor utility obtained from the fulfilled financial goals. By eliminating separate training and optimization for each new investor problem, the MetaRL model in inference mode produces near-optimal dynamic investment portfolio and goal-fulfilling strategies for a new GBWM problem within a few hundredths of a second. This delivers expected utilities that are, on average, 97.8% of the optimal expected utilities (determined via Dynamic Programming). These results are remarkably robust to capital market regime changes, even when training uses only one capital market regime. Further, the MetaRL approach can enable solving problems with larger state spaces where Dynamic Programming becomes computationally infeasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。