用强化学习优化遗传算法,提升投资组合多目标优化效果。
Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization
- 引入强化学习动态调整遗传算法参数,实现自适应进化。
- 结合灰色关联系数筛选父代,提升解集收敛性与分布均匀性。
- 在纳斯达克组合优化中实现年化夏普比1.92,适合风险偏好分析。
在现代金融市场中,决策者越来越依赖量化方法应对多目标间的复杂权衡。本文针对有约束的多目标优化问题,提出一种融合强化学习(RL)与灰色关联系数(GRC)的NSGA-II改进方法(RL-NSGA-II-GRC),用于最小化风险并最大化投资回报。该方法通过强化学习代理在线调整进化参数,基于超体积、可行性与多样性指标进行反馈;同时设计基于GRC的二元锦标赛算子,综合支配等级、拥挤距离与理想参考点接近度生成统一评分。在Kursawe和CONSTR基准测试上,相比NSGA-II分别提升约5.8%和4.4%的收敛性能,且保持良好分布的非支配解集。在纳斯达克组合应用中,生成平滑密集的有效前沿,支持识别最大夏普比率组合(年化夏普比=1.92)及不同风险偏好的效用最优组合。主要贡献包括:1)构建融合强化学习的自适应控制框架;2)设计基于GRC的综合性选择机制;3)在基准与真实金融案例中验证了更优收敛性与可解释性前沿。
原文摘要 · Abstract (English)
In modern financial markets, decision-makers increasingly rely on quantitative methods to navigate complex trade-offs among multiple, often conflicting objectives. This paper addresses constrained multi-objective optimization (MOO) with an application to portfolio optimization for minimizing risk and maximizing return. To address existing gaps, we propose a novel reinforcement learning (RL)-guided non-dominated sorting genetic algorithm II (NSGA-II) enhanced with gray relational coefficients (GRC), termed RL-NSGA-II-GRC, which combines an RL agent controller and GRC-based selection to improve convergence and diversity of Pareto fronts. The agent adapts evolutionary parameters online using metrics of hypervolume, feasibility, and diversity, while the GRC tournament operator ranks parents via a unified score considering dominance rank, crowding distance, and proximity to ideal reference. We evaluate the framework on the Kursawe and CONSTR benchmarks and a NASDAQ portfolio application. On the benchmarks, RL-NSGA-II-GRC achieves convergence improvements of about 5.8% and 4.4% over NSGA-II, while preserving well-distributed non-dominated solutions. In the portfolio application, it produces a smooth, densely populated efficient frontier supporting identification of the maximum Sharpe ratio portfolio (annualized Sharpe =1.92) and utility-optimal portfolios for different risk-aversion levels. The main contributions are three-fold: 1) we propose an RL-NSGA-II-GRC method integrating an RL agent into the evolutionary framework to adaptively control parameters via generational feedback; 2) we design a GRC-enhanced binary tournament operator providing a comprehensive indicator to guide the search toward the Pareto front; 3) we demonstrate, on benchmark MOO and a NASDAQ case study, that the method delivers improved convergence and well-populated frontiers supporting actionable insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。